Semantic enrichment
Two independent semantic layers, both optional. Without them, the graph is purely structural (AST-extracted) — still fully queryable.
Local embeddings (no API key)
astria run . --embed
Downloads a local model once (jina-embeddings-v2-base-code, ~615 MB, then offline forever) and computes vector embeddings for every node. This adds:
similar_toedges (INFERRED, cosine-scored) linking semantically related symbols across files — they flow into clustering, surprising connections, and every export- embedding-backed query recall:
querymerges semantic candidates with token matching, so conceptual questions with zero string overlap still find their symbols. When the question's identifying terms have no lexical evidence anywhere, a strong embedding match ranks like a label match instead of capping below partial word matches — the model is the only witness to what the question is about. The model loads once per process, so steady-state queries answer in ~100ms
Once embeddings exist, every run/update refreshes them incrementally (offline — the refresh never downloads), and query picks them up automatically.
Override the model cache location with ASTRIA_EMBED_CACHE_DIR — all variables in Environment variables.
LLM enrichment
Select an LLM backend explicitly with --backend or ASTRIA_LLM_BACKEND to enrich documents, papers, and images. Credentials or endpoint variables alone never activate network enrichment. --backend none or ASTRIA_LLM_BACKEND=none disables it.
| Backend | Env vars | Vision |
|---|---|---|
Anthropic Claude (claude) | ASTRIA_LLM_API_KEY | ✓ |
OpenAI-compatible (openai: OpenAI, DeepSeek, Ollama, LM Studio, custom) | ASTRIA_LLM_BASE_URL + ASTRIA_LLM_API_KEY/OPENAI_API_KEY | ✓ |
Google Gemini (gemini) | GEMINI_API_KEY or GOOGLE_API_KEY | ✓ |
ASTRIA_LLM_BACKENDselects the backend explicitlyASTRIA_LLM_MODELoverrides the model- Per-run:
astria run . --backend openai --model gpt-4o-mini - Images (png/jpg/webp/gif, ≤5 MB) go through each backend's vision API
ASTRIA_LLM_CONCURRENCYcontrols the parallel worker pool; long files are chunked and LLM output is validated
Once selected, a backend reads its configured credentials, including the generic provider environment variables listed above. Without explicit selection, the structural pipeline does not invoke an LLM. --label-communities and --deep also require explicit backend selection.
Jev judge layer
astria run . --backend openai --judge jev keeps the engine backend as the generator and layers TypeSafe's Jev on top of it. Jev is a System One decision model — typed judgments with calibrated probabilities, not a text generator — so it never writes the extraction itself; it re-judges what the engine produced:
- Gates trivial files before their first extraction with batched keep/drop judgments (≈1 request per 50 files), so empty or trivial files never cost an engine call. Gated files keep their structural extraction; the run summary reports the count ("N files gated by Jev"). Verdicts are cached per file content and judge configuration, so re-runs are reproducible and unchanged files cost no gate calls (
ASTRIA_LLM_JEV_GATE_CACHE=offrestores live re-judging). - Re-judges every extraction in one request per file: relations and node types are re-chosen from the schema allowlists (replacing the lossy clamps), and every edge gets a keep/drop existence verdict. Spurious edges are dropped; kept edges carry the judge's keep probability as a calibrated
confidence_scorein the graph — surfaced on queryEDGElines ([SEMANTIC:0.82]),explain/MCP neighbor listings, and wiki relation sections. Query traversal makes the signal load-bearing: nodes reached only through low-confidence semantic links sort after strongly-reached ones, andASTRIA_QUERY_MIN_SEMANTIC_CONFIDENCEsets a hard floor that drops weak verified edges from traversal entirely (structural and inferred edges are never filtered). - Re-ranks the suggested questions in
graph_report.mdso the most useful one leads (only on runs that rebuilt the graph).
Judge decisions are cheap and batched, count toward ASTRIA_LLM_BUDGET like every other call, and fingerprint into the extraction cache — changing the judge or its configuration invalidates cached extractions. The judge requires an explicit backend (it wraps an engine; it cannot generate extractions — --backend jev is rejected with a pointer to --judge).
Cost: measured, capped, and cached
Every response's usage block is counted across the whole run — extraction, gate, verification, community naming, and deep linking — and the run summary prints it:
LLM usage: 17 API calls, 3366 in / 1675 out tokens
ASTRIA_LLM_BUDGET(total tokens) stops extraction before the cap is exceeded; remaining files fail loudly instead of silently skipping.pipeline_runsrecords each run'sllm_input_tokens/llm_output_tokens/llm_api_calls, so spend is queryable history, not a vibe.- Extraction caches include file content plus the effective backend, endpoint, model, and prompt configuration. Matching inputs reuse output; configuration changes invalidate it. Cached and fresh output follow the same merge path.
- Failed semantic extraction leaves the core graph and successful-file manifest unadvanced, so a later update retries the work. Derived community-label and deep-link stages run after the core commit and can be retried separately on the next run.
Thematic community labels — --label-communities
astria run . --backend openai --label-communities names communities with one LLM call per changed community:
Communities labeled: 6 (--label-communities)
...
Communities labeled: 0, 6 unchanged (--label-communities) # second run: all reused
- The label ships with a one-line summary, and both are first-class data: MCP
list_communities,graph_report.md,graph.json(communitiesarray), wiki, and Obsidian all show them. - Provenance is explicit —
label_sourceisllmorhub, so a themed label is always distinguishable from the deterministic thematic/hub term that names the community otherwise. - Caching includes community membership, prompt inputs, and effective backend configuration: changed members, prompts, endpoint, or model require fresh output.
ASTRIA_LLM_COMMUNITY_MAXcaps calls per run (default 48, largest communities first); communities under 3 nodes keep deterministic names.
Deep concept links — --deep
--deep is the second extraction tier: one LLM call per file offers the file's code symbols against the graph's concept nodes and writes the meaningful matches as INFERRED edges tagged context='deep' — the cross-file concept mesh the AST cannot see.
- Results are cached against file content, offered symbols and concepts, prompts, and effective backend configuration. Changes to any of those inputs can require a fresh call.
- Source changes invalidate deep edges. Run with
--deepagain to replay valid cached links or regenerate links for changed inputs; ordinary updates do not restore them.--detail highcontinues to filter deep links as inferred evidence.
Deep concept links: 27 (--deep)