Troubleshooting
Status: Phase 1 (v0.1). Common errors and how to recover. Last verified against
mainon 2026-05-10. If you hit something not listed here, checkdocs/runbooks/for ops-level scenarios or open a Linear issue.
Postgres or Neo4j unreachable
Symptom: Server logs show
could not translate host name "postgres" or
Failed to establish connection (Neo4j). The gateway boots anyway.
Diagnosis: MetaForge degrades gracefully — without those backends it falls back to in-memory implementations.
What's lost in fallback mode:
- No persistence. Data evaporates when the gateway restarts.
- Limited Cypher.
twin.query_cypherruns against an in-memory shim; complex graph patterns may behave differently from real Neo4j. - No vector search across processes.
knowledge.*adapters drop silently if Postgres + pgvector aren't reachable; ingest returns an error.
Fix:
docker compose up -d postgres neo4j
docker compose ps # verify both report "healthy"
If you need fallback to fail loudly instead of silently, set
METAFORGE_REQUIRE_NEO4J=true and METAFORGE_REQUIRE_POSTGRES=true
in the gateway environment — the server then refuses to start
without them.
knowledge.ingest reports success but knowledge.search finds nothing
Symptom: An ingest returns chunks_indexed = N (N > 0), yet a
follow-up knowledge.search for the same content returns zero hits.
Cause: LightRAG's ainsert can fail to persist the chunk vectors
without raising — an embedding or KG-extraction error inside its
pipeline is swallowed and the call still returns normally. ingest
used to report chunks_indexed straight from the submitted chunk list,
so a silent write failure was indistinguishable from success.
Behaviour now: LightRAGKnowledgeService.ingest reads the store
back after ainsert (_count_persisted_chunks) and confirms the
chunks actually landed:
- 0 persisted while chunks were produced → ingest raises and logs
lightrag_ingest_not_persisted(withexpected_chunks/persisted_chunks), so the tool surfaces an error envelope instead of a phantom success. Retry, or check why the write failed. - Fewer persisted than produced → ingest still succeeds but logs
lightrag_ingest_partial_persistfor observability.
Diagnosis: query the store directly and check LightRAG's doc status:
SELECT count(*) FROM lightrag_vdb_chunks
WHERE workspace = 'lightrag'
AND file_path::jsonb->>'src' = '<your source_path>';
SELECT id, status, error_msg FROM lightrag_doc_status
WHERE workspace = 'lightrag' ORDER BY updated_at DESC LIMIT 10;
If the embedding worker is the culprit, the gateway logs a matching
error around the ingest — filter Loki on
scope_name = "digital_twin.knowledge.lightrag_service".
Every knowledge.search fails with invalid input syntax for type json
Symptom: all knowledge searches in a workspace error out with a
Postgres message like invalid input syntax for type json … Token "…" is invalid, regardless of query or filters.
Cause: a row in lightrag_vdb_chunks whose file_path is not the
encoded-JSON metadata blob the read paths parse. The known writer of
such rows is lightrag-hku 1.5.x, which basenames file_path on
write (keeps only the part after the last /) — pyproject.toml pins
lightrag-hku>=1.4,<1.5 for exactly this reason (MET-577). One bad row
used to make every file_path::jsonb cast in the workspace fatal.
Behaviour now: the casting queries prefilter to JSON-shaped rows
(_JSON_FILE_PATH_GUARD in lightrag_service.py), so garbage rows are
invisible rather than fatal, and the post-ingest persistence read-back
matches on chunk id without casting at all.
Diagnosis / cleanup: find (and, after inspection, remove) mangled rows:
SELECT id, left(file_path, 80) FROM lightrag_vdb_chunks
WHERE file_path !~ '^\s*[\[{"]';
Also verify the installed LightRAG version matches the pin in every
container that ingests (gateway and mcp-http sidecar — they build
from the same Dockerfile but may have been built at different times):
docker exec metaforge-gateway-1 pip show lightrag-hku | grep Version
docker exec metaforge-mcp-http-1 pip show lightrag-hku | grep Version
.mcp.json drift breaks test_mcp_json_config
Symptom: pytest tests/unit/test_mcp_json_config.py fails with:
AssertionError: assert '.venv/bin/python' == 'python'
Cause: Local edits to .mcp.json (often automatic from an IDE)
swap the canonical "command": "python" for a venv-relative path.
Fix:
git restore .mcp.json
If you get error: unable to unlink old '.mcp.json': Device or resource busy on WSL2, see the next section.
WSL2 file locks (Device or resource busy)
Symptom: git restore or git checkout of one of these files
errors out with "Device or resource busy":
.git/config.mcp.json.claude/settings.local.json
Cause: A Windows-side process (Claude Code, an IDE, an antivirus) holds the file open. The Windows handle blocks Linux from unlinking it during a git operation.
Fix: Use git show HEAD:<file> to read the canonical content,
then write it through the editor / Write tool you're already in. The
overwrite uses Linux-native fs semantics and bypasses the lock:
git show HEAD:.mcp.json > /tmp/canonical.json
# then copy the content into .mcp.json via your editor
Don't kill the holding process unless you know what it is — it's usually your active shell or IDE.
MCP error: <arg> required and must be a string
Symptom: Calling a tool from Claude Code fails with
'tool_name' is required and must be a string (or similar) for
every spec-compliant call.
Cause: This was MET-420 — the unified MCP server was dropping
the arguments payload on tools/call for the standard MCP
shape. The fix landed in PR #181.
Fix: Pull main and rebuild. If you're still seeing the error
after that, double-check your client is sending the standard shape
({name, arguments}) and not the legacy shape ({tool_id, parameters}) — the server accepts both, but mixed shapes confuse
older clients.
Dashboard 404 on /knowledge
Symptom: The dashboard renders, but /knowledge shows a 404 or
the table is permanently empty.
Cause: The gateway's knowledge adapter isn't loaded. The
dashboard route exists either way; the data behind it doesn't.
Fix: Confirm the env var:
echo $METAFORGE_ADAPTERS
# expect: cadquery,calculix,knowledge (or similar including knowledge)
If knowledge is missing, the standalone server and gateway both
skip the adapter. Set the env var before launching:
export METAFORGE_ADAPTERS=cadquery,calculix,knowledge
docker compose up gateway dashboard
You'll also need pip install -e ".[knowledge]" so the LightRAG /
asyncpg deps are present.
"Adapter X dropped silently at startup"
Symptom: cadquery.* or freecad.* tools don't show up in
tool/list even though METAFORGE_ADAPTERS includes them. No
error is logged.
Cause: The optional Python deps aren't installed. The launcher declares the manifest but skips the handler — by design, so a bare clone still boots.
Fix: Install the matching extras:
pip install -e ".[knowledge,cadquery]" # or any subset you need
Available extras: dev, knowledge, cadquery, freecad, kicad,
neo4j. Check pyproject.toml [project.optional-dependencies]
for the current list.
CLI: connection refused against the gateway
Symptom: python -m cli.forge_cli proposals exits with
Error: failed to connect to gateway: ....
Cause: No gateway is running, or it's on a different host/port from what the CLI is dialing.
Fix:
# Confirm a gateway is listening:
curl http://localhost:8000/health
# Or override the URL:
python -m cli.forge_cli --gateway-url http://gateway.local:8000 proposals
# Or set the env var:
export METAFORGE_GATEWAY_URL=http://gateway.local:8000
If you don't have a gateway anywhere, boot one:
docker compose up gateway
# or, locally (no Docker):
python -m api_gateway.server
Ingest: httpx.ReadTimeout
Symptom: python -m cli.forge_cli ingest large-file.pdf errors
out with a read timeout.
Cause: Embedding a large doc takes longer than the default 300 s.
Fix: Bump the per-request timeout:
python -m cli.forge_cli ingest large-file.pdf --timeout 1800
# or persistently:
export METAFORGE_INGEST_TIMEOUT=1800
pytest flake on first run
Symptom: First-ever pytest run on a fresh clone reports a
random unrelated failure that doesn't reproduce on the second run.
Cause: Python __pycache__ from a prior install of a different
version. Common after switching branches or pulling.
Fix:
find . -name __pycache__ -type d -prune -exec rm -rf {} +
pytest
pytest stalls near the very end (98%+), then finishes
Symptom: pytest tests/unit reaches ~98%, then sits there for
30-50 seconds with the wall clock climbing but no CPU time accruing.
/proc/<pid> shows state S, blocked on futex_wait_queue, with a
pile of threads alive. It reads exactly like a deadlock, but the run
does eventually finish.
Cause: api_gateway/server.py calls init_observability() at
module scope, so merely importing the app stands up three live OTLP
exporters aimed at OTEL_EXPORTER_OTLP_ENDPOINT, which defaults to
http://localhost:4317. With no collector listening — a test run, a
laptop without the observability stack — the batch span/metric/log
processors spend the whole of interpreter shutdown trying to flush to a
dead endpoint. Measured on the full unit suite: 202s with export on,
101s with it off, on identical code.
Fix: none needed; tests/conftest.py sets
METAFORGE_OTEL_EXPORT=off before any app import. If you see this
stall, check that line still exists —
test_export_is_off_for_this_run fails loudly if it is removed.
To profile telemetry deliberately, set METAFORGE_OTEL_EXPORT=on and
expect the suite to take about twice as long.
Telemetry env switches
Two switches, and the difference matters:
| Variable | Effect | Use it when |
|---|---|---|
METAFORGE_OTEL_EXPORT=off | Builds no exporters. Tracing stays fully functional, so instrumentation still records spans — anything that installs its own span processor keeps working. | You have no collector: tests, local dev, CI |
OTEL_SDK_DISABLED=true | The OpenTelemetry-standard kill switch. Turns the SDK off entirely, so code records nothing. | You want no telemetry at all |
Reaching for OTEL_SDK_DISABLED when you only meant "don't export" is
the trap: it also makes the SDK hand out no-op tracers, which silently
breaks any test that asserts on span attributes.
The MCP sidecar is serving old code (FORGE-411)
Symptoms read like a smaller deployment, not like staleness: an older
protocol version in the initialize result, tools missing from
tools/list, a prompt that does not exist. The first time this happened,
four bugs were filed against the run and two of their findings were not
real.
The cause is that mcp-http and gateway pick up a new checkout
differently:
| Service | How it loads code | A git checkout on the host |
|---|---|---|
gateway | uvicorn's reloader supervises it (--reload) | picked up within seconds |
mcp-http | calls uvicorn.Server.serve() in the bootstrap's own event loop, so there is no reloader to supervise it (MET-477 G3) | ignored until the container restarts |
So the sidecar needs restarting on deploy. It is not a rebuild:
docker compose up -d --force-recreate mcp-http
mcp-http runs the CI-published gateway image rather than a local build,
so after a merge the sequence is docker compose pull mcp-http then the
command above. Do not docker compose build mcp-http — the service
has no build: stanza on purpose, because the same Dockerfile built two
ways can diverge without anything saying so.
Ask instead of inferring
health.check answers it directly now — it is a tool, so any harness can
call it with no shell:
"code": {
"build_sha": "97020586ec2b",
"source_sha": "4f1ac0b77e91",
"stale": true,
"result": "stale",
"reloads": false
}
A stale process reports "status": "stale" rather than "healthy": it
answers every request correctly, for code nobody is reading.
Two caveats worth knowing before you trust it:
result: "unknown"means the image was built without--build-arg METAFORGE_BUILD_SHA=$(git rev-parse HEAD), so the check cannot run at all. CI passes it; a hand-built image may not.reloads: true(the gateway) never reports stale, because for a supervised process a difference between the two SHAs is expected and transient.
Prometheus carries the same thing as
metaforge_mcp_code_version_check_total{result="stale"}, which is what
the McpRunningStaleCode alert fires on — the point being to be told
without looking.
docker compose fails: required variable <NAME> is missing a value
Every docker compose command fails, including ones that have nothing to do
with the named service:
error while interpolating services.postgres.environment.POSTGRES_PASSWORD:
required variable POSTGRES_PASSWORD is missing a value
This is deliberate (FORGE-556, FORGE-558). These credentials used to have
working defaults — metaforge, minioadmin, temporal — committed to this
public repository since the first compose commit. Every MetaForge install
therefore shared passwords that any reader of the repository already knew, on
services that publish host ports by default:
| Variable | Old default | Service | Published on |
|---|---|---|---|
POSTGRES_PASSWORD | metaforge | Postgres | 5432 |
NEO4J_PASSWORD | metaforge | Neo4j | 7474, 7687 |
MINIO_SECRET_KEY | minioadmin | MinIO | 9000, 9001 |
TEMPORAL_POSTGRES_PASSWORD | temporal | Temporal's database | not published |
GRAFANA_PASSWORD | metaforge | Grafana | 3001, --profile observability |
The defaults are gone, and compose refuses rather than picking one for you.
Identifiers that are not secrets — POSTGRES_USER, POSTGRES_DB,
NEO4J_USER, MINIO_ACCESS_KEY — keep their defaults.
Compose interpolates the whole file before it decides which services to start,
so the error appears even for docker compose up gateway, where the named
service is not involved. That is a property of compose, not a bug here.
Fix it in either direction:
scripts/onboarding.sh # fills every blank, in both --usage and --develop
# or, by hand, for each one the error names:
for v in POSTGRES_PASSWORD NEO4J_PASSWORD MINIO_SECRET_KEY \
TEMPORAL_POSTGRES_PASSWORD GRAFANA_PASSWORD; do
echo "$v=$(openssl rand -base64 24)" >> .env
done
Fix them one at a time and compose will name the next one, since interpolation stops at the first failure.
If you copied .env.example and skipped onboarding, you will hit this: the file
ships these keys blank on purpose. onboarding.sh generates secrets with
set_if_blank, which returns early for any key that already has a value — so
the old shipped defaults silently guaranteed the generator never fired, and
every install kept the published passwords.
Already running services that used the old defaults? Setting the variables is
not enough. Postgres, Neo4j, MinIO and Grafana all seed their credentials when
they first provision their data directory, and those live in named volumes that
survive docker compose down. An existing stack keeps the published passwords
until each service is rotated in place, or its volume is deleted and the stack
recreated from scratch (which destroys the data in it).
For Grafana specifically: GF_SECURITY_ADMIN_PASSWORD seeds the admin user when Grafana first
provisions its database; on later starts an existing admin keeps the password it
already has. Since grafana-data is a named volume that survives
docker compose down, an instance created before this change is still on the
published default until you rotate it explicitly:
# Grafana 9+ (the `latest` image). Older images use `grafana-cli` instead.
docker compose exec grafana grafana cli admin reset-admin-password '<new password>'
Confirm which one your image has with
docker compose exec grafana sh -c 'command -v grafana grafana-cli'.
When to escalate
If the issue is:
- a missing capability — file a Linear issue under the appropriate
epic (see
roadmap.md). - a regression — open a Linear ticket and tag it
regression; the bug-hunter agent (/bug-hunt) can help triage. - an ops-level outage — see
docs/runbooks/for stack-specific runbooks (gateway-down.md,neo4j-unreachable.md,kafka-consumer-stopped.md).