Skip to main content

Troubleshooting

Status: Phase 1 (v0.1). Common errors and how to recover. Last verified against main on 2026-05-10. If you hit something not listed here, check docs/runbooks/ for ops-level scenarios or open a Linear issue.

Postgres or Neo4j unreachable​

Symptom: Server logs show could not translate host name "postgres" or Failed to establish connection (Neo4j). The gateway boots anyway.

Diagnosis: MetaForge degrades gracefully — without those backends it falls back to in-memory implementations.

What's lost in fallback mode:

  • No persistence. Data evaporates when the gateway restarts.
  • Limited Cypher. twin.query_cypher runs against an in-memory shim; complex graph patterns may behave differently from real Neo4j.
  • No vector search across processes. knowledge.* adapters drop silently if Postgres + pgvector aren't reachable; ingest returns an error.

Fix:

Terminal
docker compose up -d postgres neo4j
docker compose ps # verify both report "healthy"

If you need fallback to fail loudly instead of silently, set METAFORGE_REQUIRE_NEO4J=true and METAFORGE_REQUIRE_POSTGRES=true in the gateway environment — the server then refuses to start without them.

knowledge.ingest reports success but knowledge.search finds nothing​

Symptom: An ingest returns chunks_indexed = N (N > 0), yet a follow-up knowledge.search for the same content returns zero hits.

Cause: LightRAG's ainsert can fail to persist the chunk vectors without raising — an embedding or KG-extraction error inside its pipeline is swallowed and the call still returns normally. ingest used to report chunks_indexed straight from the submitted chunk list, so a silent write failure was indistinguishable from success.

Behaviour now: LightRAGKnowledgeService.ingest reads the store back after ainsert (_count_persisted_chunks) and confirms the chunks actually landed:

  • 0 persisted while chunks were produced → ingest raises and logs lightrag_ingest_not_persisted (with expected_chunks / persisted_chunks), so the tool surfaces an error envelope instead of a phantom success. Retry, or check why the write failed.
  • Fewer persisted than produced → ingest still succeeds but logs lightrag_ingest_partial_persist for observability.

Diagnosis: query the store directly and check LightRAG's doc status:

SQL
SELECT count(*) FROM lightrag_vdb_chunks
WHERE workspace = 'lightrag'
AND file_path::jsonb->>'src' = '<your source_path>';

SELECT id, status, error_msg FROM lightrag_doc_status
WHERE workspace = 'lightrag' ORDER BY updated_at DESC LIMIT 10;

If the embedding worker is the culprit, the gateway logs a matching error around the ingest — filter Loki on scope_name = "digital_twin.knowledge.lightrag_service".

Every knowledge.search fails with invalid input syntax for type json​

Symptom: all knowledge searches in a workspace error out with a Postgres message like invalid input syntax for type json … Token "…" is invalid, regardless of query or filters.

Cause: a row in lightrag_vdb_chunks whose file_path is not the encoded-JSON metadata blob the read paths parse. The known writer of such rows is lightrag-hku 1.5.x, which basenames file_path on write (keeps only the part after the last /) — pyproject.toml pins lightrag-hku>=1.4,<1.5 for exactly this reason (MET-577). One bad row used to make every file_path::jsonb cast in the workspace fatal.

Behaviour now: the casting queries prefilter to JSON-shaped rows (_JSON_FILE_PATH_GUARD in lightrag_service.py), so garbage rows are invisible rather than fatal, and the post-ingest persistence read-back matches on chunk id without casting at all.

Diagnosis / cleanup: find (and, after inspection, remove) mangled rows:

SQL
SELECT id, left(file_path, 80) FROM lightrag_vdb_chunks
WHERE file_path !~ '^\s*[\[{"]';

Also verify the installed LightRAG version matches the pin in every container that ingests (gateway and mcp-http sidecar — they build from the same Dockerfile but may have been built at different times):

Terminal
docker exec metaforge-gateway-1 pip show lightrag-hku | grep Version
docker exec metaforge-mcp-http-1 pip show lightrag-hku | grep Version

.mcp.json drift breaks test_mcp_json_config​

Symptom: pytest tests/unit/test_mcp_json_config.py fails with:

Code
AssertionError: assert '.venv/bin/python' == 'python'

Cause: Local edits to .mcp.json (often automatic from an IDE) swap the canonical "command": "python" for a venv-relative path.

Fix:

Terminal
git restore .mcp.json

If you get error: unable to unlink old '.mcp.json': Device or resource busy on WSL2, see the next section.

WSL2 file locks (Device or resource busy)​

Symptom: git restore or git checkout of one of these files errors out with "Device or resource busy":

  • .git/config
  • .mcp.json
  • .claude/settings.local.json

Cause: A Windows-side process (Claude Code, an IDE, an antivirus) holds the file open. The Windows handle blocks Linux from unlinking it during a git operation.

Fix: Use git show HEAD:<file> to read the canonical content, then write it through the editor / Write tool you're already in. The overwrite uses Linux-native fs semantics and bypasses the lock:

Terminal
git show HEAD:.mcp.json > /tmp/canonical.json
# then copy the content into .mcp.json via your editor

Don't kill the holding process unless you know what it is — it's usually your active shell or IDE.

MCP error: <arg> required and must be a string​

Symptom: Calling a tool from Claude Code fails with 'tool_name' is required and must be a string (or similar) for every spec-compliant call.

Cause: This was MET-420 — the unified MCP server was dropping the arguments payload on tools/call for the standard MCP shape. The fix landed in PR #181.

Fix: Pull main and rebuild. If you're still seeing the error after that, double-check your client is sending the standard shape ({name, arguments}) and not the legacy shape ({tool_id, parameters}) — the server accepts both, but mixed shapes confuse older clients.

Dashboard 404 on /knowledge​

Symptom: The dashboard renders, but /knowledge shows a 404 or the table is permanently empty.

Cause: The gateway's knowledge adapter isn't loaded. The dashboard route exists either way; the data behind it doesn't.

Fix: Confirm the env var:

Terminal
echo $METAFORGE_ADAPTERS
# expect: cadquery,calculix,knowledge (or similar including knowledge)

If knowledge is missing, the standalone server and gateway both skip the adapter. Set the env var before launching:

Terminal
export METAFORGE_ADAPTERS=cadquery,calculix,knowledge
docker compose up gateway dashboard

You'll also need pip install -e ".[knowledge]" so the LightRAG / asyncpg deps are present.

"Adapter X dropped silently at startup"​

Symptom: cadquery.* or freecad.* tools don't show up in tool/list even though METAFORGE_ADAPTERS includes them. No error is logged.

Cause: The optional Python deps aren't installed. The launcher declares the manifest but skips the handler — by design, so a bare clone still boots.

Fix: Install the matching extras:

Terminal
pip install -e ".[knowledge,cadquery]" # or any subset you need

Available extras: dev, knowledge, cadquery, freecad, kicad, neo4j. Check pyproject.toml [project.optional-dependencies] for the current list.

CLI: connection refused against the gateway​

Symptom: python -m cli.forge_cli proposals exits with Error: failed to connect to gateway: ....

Cause: No gateway is running, or it's on a different host/port from what the CLI is dialing.

Fix:

Terminal
# Confirm a gateway is listening:
curl http://localhost:8000/health

# Or override the URL:
python -m cli.forge_cli --gateway-url http://gateway.local:8000 proposals
# Or set the env var:
export METAFORGE_GATEWAY_URL=http://gateway.local:8000

If you don't have a gateway anywhere, boot one:

Terminal
docker compose up gateway
# or, locally (no Docker):
python -m api_gateway.server

Ingest: httpx.ReadTimeout​

Symptom: python -m cli.forge_cli ingest large-file.pdf errors out with a read timeout.

Cause: Embedding a large doc takes longer than the default 300 s.

Fix: Bump the per-request timeout:

Terminal
python -m cli.forge_cli ingest large-file.pdf --timeout 1800
# or persistently:
export METAFORGE_INGEST_TIMEOUT=1800

pytest flake on first run​

Symptom: First-ever pytest run on a fresh clone reports a random unrelated failure that doesn't reproduce on the second run.

Cause: Python __pycache__ from a prior install of a different version. Common after switching branches or pulling.

Fix:

Terminal
find . -name __pycache__ -type d -prune -exec rm -rf {} +
pytest

pytest stalls near the very end (98%+), then finishes​

Symptom: pytest tests/unit reaches ~98%, then sits there for 30-50 seconds with the wall clock climbing but no CPU time accruing. /proc/<pid> shows state S, blocked on futex_wait_queue, with a pile of threads alive. It reads exactly like a deadlock, but the run does eventually finish.

Cause: api_gateway/server.py calls init_observability() at module scope, so merely importing the app stands up three live OTLP exporters aimed at OTEL_EXPORTER_OTLP_ENDPOINT, which defaults to http://localhost:4317. With no collector listening — a test run, a laptop without the observability stack — the batch span/metric/log processors spend the whole of interpreter shutdown trying to flush to a dead endpoint. Measured on the full unit suite: 202s with export on, 101s with it off, on identical code.

Fix: none needed; tests/conftest.py sets METAFORGE_OTEL_EXPORT=off before any app import. If you see this stall, check that line still exists — test_export_is_off_for_this_run fails loudly if it is removed.

To profile telemetry deliberately, set METAFORGE_OTEL_EXPORT=on and expect the suite to take about twice as long.

Telemetry env switches​

Two switches, and the difference matters:

VariableEffectUse it when
METAFORGE_OTEL_EXPORT=offBuilds no exporters. Tracing stays fully functional, so instrumentation still records spans — anything that installs its own span processor keeps working.You have no collector: tests, local dev, CI
OTEL_SDK_DISABLED=trueThe OpenTelemetry-standard kill switch. Turns the SDK off entirely, so code records nothing.You want no telemetry at all

Reaching for OTEL_SDK_DISABLED when you only meant "don't export" is the trap: it also makes the SDK hand out no-op tracers, which silently breaks any test that asserts on span attributes.

The MCP sidecar is serving old code (FORGE-411)​

Symptoms read like a smaller deployment, not like staleness: an older protocol version in the initialize result, tools missing from tools/list, a prompt that does not exist. The first time this happened, four bugs were filed against the run and two of their findings were not real.

The cause is that mcp-http and gateway pick up a new checkout differently:

ServiceHow it loads codeA git checkout on the host
gatewayuvicorn's reloader supervises it (--reload)picked up within seconds
mcp-httpcalls uvicorn.Server.serve() in the bootstrap's own event loop, so there is no reloader to supervise it (MET-477 G3)ignored until the container restarts

So the sidecar needs restarting on deploy. It is not a rebuild:

Terminal
docker compose up -d --force-recreate mcp-http

mcp-http runs the CI-published gateway image rather than a local build, so after a merge the sequence is docker compose pull mcp-http then the command above. Do not docker compose build mcp-http — the service has no build: stanza on purpose, because the same Dockerfile built two ways can diverge without anything saying so.

Ask instead of inferring​

health.check answers it directly now — it is a tool, so any harness can call it with no shell:

JSON
"code": {
"build_sha": "97020586ec2b",
"source_sha": "4f1ac0b77e91",
"stale": true,
"result": "stale",
"reloads": false
}

A stale process reports "status": "stale" rather than "healthy": it answers every request correctly, for code nobody is reading.

Two caveats worth knowing before you trust it:

  • result: "unknown" means the image was built without --build-arg METAFORGE_BUILD_SHA=$(git rev-parse HEAD), so the check cannot run at all. CI passes it; a hand-built image may not.
  • reloads: true (the gateway) never reports stale, because for a supervised process a difference between the two SHAs is expected and transient.

Prometheus carries the same thing as metaforge_mcp_code_version_check_total{result="stale"}, which is what the McpRunningStaleCode alert fires on — the point being to be told without looking.

docker compose fails: required variable <NAME> is missing a value​

Every docker compose command fails, including ones that have nothing to do with the named service:

Code
error while interpolating services.postgres.environment.POSTGRES_PASSWORD:
required variable POSTGRES_PASSWORD is missing a value

This is deliberate (FORGE-556, FORGE-558). These credentials used to have working defaults — metaforge, minioadmin, temporal — committed to this public repository since the first compose commit. Every MetaForge install therefore shared passwords that any reader of the repository already knew, on services that publish host ports by default:

VariableOld defaultServicePublished on
POSTGRES_PASSWORDmetaforgePostgres5432
NEO4J_PASSWORDmetaforgeNeo4j7474, 7687
MINIO_SECRET_KEYminioadminMinIO9000, 9001
TEMPORAL_POSTGRES_PASSWORDtemporalTemporal's databasenot published
GRAFANA_PASSWORDmetaforgeGrafana3001, --profile observability

The defaults are gone, and compose refuses rather than picking one for you. Identifiers that are not secrets — POSTGRES_USER, POSTGRES_DB, NEO4J_USER, MINIO_ACCESS_KEY — keep their defaults.

Compose interpolates the whole file before it decides which services to start, so the error appears even for docker compose up gateway, where the named service is not involved. That is a property of compose, not a bug here.

Fix it in either direction:

Terminal
scripts/onboarding.sh # fills every blank, in both --usage and --develop
# or, by hand, for each one the error names:
for v in POSTGRES_PASSWORD NEO4J_PASSWORD MINIO_SECRET_KEY \
TEMPORAL_POSTGRES_PASSWORD GRAFANA_PASSWORD; do
echo "$v=$(openssl rand -base64 24)" >> .env
done

Fix them one at a time and compose will name the next one, since interpolation stops at the first failure.

If you copied .env.example and skipped onboarding, you will hit this: the file ships these keys blank on purpose. onboarding.sh generates secrets with set_if_blank, which returns early for any key that already has a value — so the old shipped defaults silently guaranteed the generator never fired, and every install kept the published passwords.

Already running services that used the old defaults? Setting the variables is not enough. Postgres, Neo4j, MinIO and Grafana all seed their credentials when they first provision their data directory, and those live in named volumes that survive docker compose down. An existing stack keeps the published passwords until each service is rotated in place, or its volume is deleted and the stack recreated from scratch (which destroys the data in it).

For Grafana specifically: GF_SECURITY_ADMIN_PASSWORD seeds the admin user when Grafana first provisions its database; on later starts an existing admin keeps the password it already has. Since grafana-data is a named volume that survives docker compose down, an instance created before this change is still on the published default until you rotate it explicitly:

Terminal
# Grafana 9+ (the `latest` image). Older images use `grafana-cli` instead.
docker compose exec grafana grafana cli admin reset-admin-password '<new password>'

Confirm which one your image has with docker compose exec grafana sh -c 'command -v grafana grafana-cli'.

When to escalate​

If the issue is:

  • a missing capability — file a Linear issue under the appropriate epic (see roadmap.md).
  • a regression — open a Linear ticket and tag it regression; the bug-hunter agent (/bug-hunt) can help triage.
  • an ops-level outage — see docs/runbooks/ for stack-specific runbooks (gateway-down.md, neo4j-unreachable.md, kafka-consumer-stopped.md).
MetaForge documentationGateway schema