Repository guide · architecture · v2
simonw / llm · Architecture
The repository is a Python project of 116 analyzed files organized into docs, llm, and tests directories. Bounded static paths start at llm/__main__.py, reach llm/cli.py:cli, and end at unresolved boundaries reached through load_plugins in llm/plugins.py. Of 9,698 static relations, 7,374 are external or unresolved — a bounded static picture, not a runtime account.
- Original repository
- simonw/llm
- License
- Apache-2.0 · open-source license
- Analyzed revision · last verified
- 1df47ddcac20d58726a993949da8ef84f4081085 ·
Newer revision observed; the code this guide cites is unchanged. The default branch moved to 764dc386c58b, checked 2026-10-03. A revision diff of 32 changed files found no change in the code regions this guide cites. The guide still describes revision 1df47ddcac20; repository-wide counts (files, calls and so on) refer to that revision.
simonw/llm architecture: a bounded static reconstruction
Terms used on this page
- Entrypoint record
- A file the analyzer marks as a place where execution can start. When it carries a __main__ guard excerpt, the file contains an if __name__ == "__main__": block and the excerpt shows what that block calls; without an excerpt, the file was flagged by its name only.
- Bounded static execution path
- A call chain reconstructed from the source code without running it, stopped after a fixed number of steps. It shows how far the code can be followed on paper, not what happens at run time.
- Unresolved boundary
- Where a static path stops because the next call goes into an external library or cannot be resolved without running the code. It marks the edge of this analysis, not a defect in the repository.
- Static relations
- Calls and imports found in the source. "External or unresolved" relations point outside the analyzed files.
- Module role
- The analyzer's label for a file, inferred from how many calls go in and out (for example core, entry/orchestration, leaf). It describes a position in the call graph, not the authors' design.
- Test-like files
- Files whose names or locations look like tests. This analysis counts them; it does not run them.
- Observed · inferred · author-claimed · unresolved
- How each statement is supported: read directly from the analyzed files; derived by the analyzer from them; stated by the repository's authors (metadata, README); or not established by this analysis.
- Verified
- Two uses on these pages. In an analyzer record ("verified provenance", "a verified entrypoint") it means the record was read directly from the analyzed files, which these guides call observed; for an entrypoint, the __main__ guard text is present. Since the analyzer fix of 2026-10-04, a file that only has an entrypoint-like name such as cli.py or main.py is recorded as inferred; guides analyzed before that still call such a file verified, and it is still only a guess about how the file is used. It does not mean the code was run or tested. "Last verified" is the date the guide last passed the lab's checks against the analysis record of the stated revision; it is not a review of the repository itself.
Shape at a glance
The analyzer's summary describes a Python project of 116 analyzed files organized around three directory subsystems: docs, llm, and tests. The inventory counts 11 root files, 6 in .github, 35 in docs, 20 in llm, and 44 in tests, with 44 test-like files noted overall.
The llm package holds the paths the role records name as active code: cli.py, plugins.py, models.py, logs.py, migrations.py, utils.py, parts.py, and default_plugins/openai_models.py. At the root, README.md, pyproject.toml, and LICENSE are flagged as project-level important, while llm/__main__.py and llm/cli.py are flagged as likely entrypoints. The only parsed dependency records come from docs/requirements.txt (sphinx and related tooling), so docs carries the documentation build.
Entry points and control flow
Two entrypoints were detected: llm/__main__.py, whose __main__ guard calls cli(), and llm/cli.py, flagged as likely executable by filename heuristic. The bounded path from llm/__main__.py is a single local step to llm/cli.py:cli, terminating as a leaf with no cycle and no truncation.
From llm/cli.py, one path ends immediately at warnings.simplefilter. Eight paths descend into llm/plugins.py [load_plugins] and stop at external or unresolved targets: hasattr, pm.load_setuptools_entrypoints, LLM_LOAD_PLUGINS.split, name.strip, metadata.distribution, entry_point.load, pm.register, and pm._plugin_distinfo.append. The reconstruction reports 13 bounded execution paths, 2 entrypoints, and 2059 resolved call relations; this packet includes ten of those paths. Read statically, llm/cli.py is the first module reached, and load_plugins is where eight of the nine traced paths from llm/cli.py leave the resolved call graph; the ninth ends at warnings.simplefilter without passing load_plugins.
Core modules and their roles
These role labels are inferred by the analyzer from static resolved call degree and entrypoint membership; they describe call-graph position, not design intent. llm/__init__.py is the most connected module — 754 incoming, 48 outgoing — and llm/parts.py follows with 470 incoming, 22 outgoing; both are labeled service/core candidates. llm/cli.py carries the entry/orchestration label (105 in, 245 out), matching its entrypoint status. Among library modules, llm/models.py (130 in / 96 out), llm/logs.py (123 in / 64 out), llm/utils.py (102 in / 12 out), and llm/default_plugins/openai_models.py (89 in / 130 out) are labeled service/core candidates. llm/plugins.py sits at the edge — 41 in, 0 out, labeled leaf/data-boundary — with llm/migrations.py nearby (38 in, 1 out). The labels track call volume: tests/test_parts.py gets the orchestration label with 422 outgoing calls, and tests/test_logs_store.py is labeled service/core.
Where static paths stop
Plugin loading is the dominant boundary: from load_plugins, the traced paths reach pm.load_setuptools_entrypoints, entry_point.load, pm.register, and metadata.distribution, all external or unresolved; the analyzer's boundary list also names importlib.import_module, though no exec path terminates there. Plain-code stops appear too: llm/cli.py ends at warnings.simplefilter, and load_plugins touches hasattr, LLM_LOAD_PLUGINS.split, and name.strip; the boundary list additionally names sys.stderr.write without an exec record. The scale is large — 7,374 of 9,698 relations are external_or_unresolved against 2,324 local, and the analyzer keeps textual targets instead of inventing local destinations. In practice, anything behind plugin loading — installed entry points, hook registration — lies outside this reconstruction, because all eight traced paths into load_plugins end at an unresolved_boundary; dynamic imports are separately recorded by the analyzer as not resolved.
Dependencies between parts
By static degree, llm/__init__.py is the package's aggregation point (754 incoming calls), and llm/cli.py is a heavy consumer (105 in, 245 out) drawing on that hub. llm/parts.py and llm/models.py carry much of the remaining local call mass, while llm/plugins.py is a leaf — 41 incoming, 0 outgoing. All seven parsed dependency records are labeled scope 'runtime' yet source from docs/requirements.txt (sphinx and documentation tooling), so they describe the docs build toolchain. Because pyproject.toml dependency tables were not parsed, the package's own runtime dependencies are not established by this analysis. Unresolved edges dominate the graph: 7,374 external_or_unresolved versus 2,324 local, and dynamic imports are explicitly not resolved.
What static analysis cannot show
This is a static picture: reflection, runtime dependency injection, dynamic imports, monkey-patching, generated code, framework runtime wiring, and dynamic dispatch are not resolved by the analyzer. How plugin loading, model resolution, and logging interact at runtime is not established by this packet. Install/run/build inference is partial and installation commands are not verified, so nothing here is a runnable recipe. Tests appear as 44 test-like files and as call-graph nodes, but the packet cannot show whether they pass or how the code behaves when executed. README-derived claims in this pipeline are line-based extracts that may capture code lines, so they stand as author statements at most.
A reading order for the architecture
Start with README.md and pyproject.toml for framing, then follow execution: llm/__main__.py, llm/cli.py:cli, and llm/plugins.py [load_plugins]. The analyzer's teaching record recommends the same shape — manifests and important symbols first, then the bounded path, then unresolved boundaries. For the core, read llm/__init__.py as the hub, then llm/parts.py and llm/models.py. Close with tests/test_parts.py, whose 422 outgoing calls make it the labeled orchestration candidate, before changing any code.
What this analysis could not establish
The writing model's own notes on the analysis record. Identifiers such as lim_3, exec_* or claim_… name records of that analysis; the claims table cites the same records.
- All control-flow statements describe bounded static paths; runtime wiring, configuration, and dynamic imports are not established (lim_1).
- The package's own runtime dependencies are unknown: only docs/requirements.txt records were parsed, and pyproject.toml dependency tables were skipped (lim_3).
- Role labels are call-count-derived and mark some test files as service/core candidates (role_30, role_35); they are not design statements.
- The reconstruction reports 13 bounded execution paths; this packet includes ten of them, so three paths are not described here.
- Install/run/build behavior is only partially inferred and installation commands are not verified (lim_2).
Claims and evidence — 35 claims, 35 supported by an independent verifier
Every substantive statement above is a claim bound to grounding IDs of the analyzed revision. Global grounding IDs are namespaced by the analysis run.
| Claim | Epistemic status | Verifier | Grounding (global IDs) |
|---|---|---|---|
| c1 The analyzer's summary describes a Python project of 116 analyzed files organized around three top-level directory subsystems: docs, llm, and tests. | inferred | supported | analysis_f20a9f60553ce57f:summary_repository, analysis_f20a9f60553ce57f:subsys_1, analysis_f20a9f60553ce57f:subsys_2, analysis_f20a9f60553ce57f:subsys_3 |
| c2 The inventory counts 11 root files, 6 in .github, 35 in docs, 20 in llm, and 44 in tests, and the maturity signals note 44 test-like files. | inferred | supported | analysis_f20a9f60553ce57f:summary_repository |
| c3 The llm package holds the modules the role records name: cli.py, plugins.py, models.py, logs.py, migrations.py, utils.py, parts.py, and default_plugins/openai_models.py. | inferred | supported | analysis_f20a9f60553ce57f:role_2, analysis_f20a9f60553ce57f:role_4, analysis_f20a9f60553ce57f:role_5, analysis_f20a9f60553ce57f:role_9, analysis_f20a9f60553ce57f:role_10, analysis_f20a9f60553ce57f:role_11, analysis_f20a9f60553ce57f:role_12, analysis_f20a9f60553ce57f:role_13, analysis_f20a9f60553ce57f:role_15 |
| c4 README.md, pyproject.toml, and LICENSE are flagged as project-level important files; llm/__main__.py and llm/cli.py are flagged as likely entrypoints. | observed | supported | analysis_f20a9f60553ce57f:important_1, analysis_f20a9f60553ce57f:important_3, analysis_f20a9f60553ce57f:important_4, analysis_f20a9f60553ce57f:important_5, analysis_f20a9f60553ce57f:important_6 |
| c5 The only parsed dependency records are seven docs/requirements.txt entries: sphinx, furo, sphinx-autobuild, sphinx-copybutton, sphinx-markdown-builder, myst-parser, and cogapp. | observed | supported | analysis_f20a9f60553ce57f:dep_1, analysis_f20a9f60553ce57f:dep_2, analysis_f20a9f60553ce57f:dep_3, analysis_f20a9f60553ce57f:dep_4, analysis_f20a9f60553ce57f:dep_5, analysis_f20a9f60553ce57f:dep_6, analysis_f20a9f60553ce57f:dep_7 |
| c6 Two entrypoints were detected: llm/__main__.py, whose __main__ guard calls cli(), and llm/cli.py, flagged as likely executable by filename heuristic. | observed | supported | analysis_f20a9f60553ce57f:py_entry_8d052da6b4aa, analysis_f20a9f60553ce57f:entry_1, analysis_f20a9f60553ce57f:important_5, analysis_f20a9f60553ce57f:important_6 |
| c7 The bounded path from llm/__main__.py has one local step, llm/__main__.py -> llm/cli.py:cli, and ends with terminal_reason 'leaf'; no cycle, not truncated. | inferred | supported | analysis_f20a9f60553ce57f:exec_1883828468e7 |
| c8 One bounded path from llm/cli.py ends immediately at warnings.simplefilter with terminal_reason 'unresolved_boundary'. | inferred | supported | analysis_f20a9f60553ce57f:exec_837f2dee781c |
| c9 Eight bounded paths run llm/cli.py -> llm/plugins.py [load_plugins] and end at external or unresolved targets: hasattr, pm.load_setuptools_entrypoints, LLM_LOAD_PLUGINS.split, name.strip, metadata.distribution, entry_point.load, pm.register, and pm._plugin_distinfo.append. | inferred | supported | analysis_f20a9f60553ce57f:exec_bbf7b5170507, analysis_f20a9f60553ce57f:exec_c714034a13f9, analysis_f20a9f60553ce57f:exec_cf7f251ec694, analysis_f20a9f60553ce57f:exec_25ee09cebebb, analysis_f20a9f60553ce57f:exec_a02b71c30924, analysis_f20a9f60553ce57f:exec_8eb173511297, analysis_f20a9f60553ce57f:exec_bfffaa01eacb, analysis_f20a9f60553ce57f:exec_6e97406cf140 |
| c10 The reconstruction reports 2 entrypoints, 2059 resolved static call relations, and 13 bounded execution paths; this packet includes ten of those paths. | inferred | supported | analysis_f20a9f60553ce57f:architecture_reconstruction, analysis_f20a9f60553ce57f:exec_1883828468e7 |
| c11 Read statically, llm/cli.py is the first module reached from the package entrypoint, and load_plugins is where eight of the nine traced paths from llm/cli.py leave the resolved call graph; the ninth ends at warnings.simplefilter without passing load_plugins. | inferred | supported | analysis_f20a9f60553ce57f:exec_1883828468e7, analysis_f20a9f60553ce57f:exec_c714034a13f9, analysis_f20a9f60553ce57f:exec_837f2dee781c |
| c12 Role labels are inferred by the analyzer from static resolved call degree and entrypoint membership; they describe call-graph position, not design intent. | inferred | supported | analysis_f20a9f60553ce57f:role_2, analysis_f20a9f60553ce57f:role_4 |
| c13 llm/__init__.py has the highest static call degree: 754 incoming and 48 outgoing calls, labeled service/core candidate. | inferred | supported | analysis_f20a9f60553ce57f:role_2 |
| c14 llm/parts.py follows: 470 incoming and 22 outgoing calls, labeled service/core candidate. | inferred | supported | analysis_f20a9f60553ce57f:role_12 |
| c15 llm/cli.py is labeled entry/orchestration candidate: 105 incoming and 245 outgoing calls. | inferred | supported | analysis_f20a9f60553ce57f:role_4 |
| c16 llm/models.py (130 in / 96 out), llm/logs.py (123 in / 64 out), llm/utils.py (102 in / 12 out), and llm/default_plugins/openai_models.py (89 in / 130 out) are labeled service/core candidates. | inferred | supported | analysis_f20a9f60553ce57f:role_11, analysis_f20a9f60553ce57f:role_9, analysis_f20a9f60553ce57f:role_15, analysis_f20a9f60553ce57f:role_5 |
| c17 llm/plugins.py is labeled leaf/data-boundary candidate with 41 incoming and 0 outgoing calls; llm/migrations.py shows 38 incoming and 1 outgoing. | inferred | supported | analysis_f20a9f60553ce57f:role_13, analysis_f20a9f60553ce57f:role_10 |
| c18 The labels track call volume: tests/test_parts.py draws the orchestration label with 422 outgoing calls, and tests/test_logs_store.py is labeled service/core candidate. | inferred | supported | analysis_f20a9f60553ce57f:role_35, analysis_f20a9f60553ce57f:role_30 |
| c19 Plugin loading is the dominant boundary: from load_plugins the traced paths reach pm.load_setuptools_entrypoints, entry_point.load, pm.register, and metadata.distribution, all external or unresolved; the analyzer's boundary list also names importlib.import_module, though no exec path terminates there. | inferred | supported | analysis_f20a9f60553ce57f:exec_c714034a13f9, analysis_f20a9f60553ce57f:exec_8eb173511297, analysis_f20a9f60553ce57f:exec_bfffaa01eacb, analysis_f20a9f60553ce57f:exec_a02b71c30924 |
| c20 Plain-code stops also appear: llm/cli.py ends at warnings.simplefilter, and load_plugins touches hasattr, LLM_LOAD_PLUGINS.split, and name.strip; the boundary list additionally names sys.stderr.write without an exec record. | inferred | supported | analysis_f20a9f60553ce57f:exec_837f2dee781c, analysis_f20a9f60553ce57f:exec_bbf7b5170507, analysis_f20a9f60553ce57f:exec_cf7f251ec694, analysis_f20a9f60553ce57f:exec_25ee09cebebb |
| c21 Of 9,698 static relations, 7,374 are external_or_unresolved against 2,324 local; the analyzer preserves textual targets rather than fabricating local destinations. | inferred | supported | analysis_f20a9f60553ce57f:relation_counts |
| c22 Anything reached through plugin loading — installed entry points, hook registration — lies outside the resolved graph, since all eight traced paths into load_plugins terminate at unresolved_boundary. | inferred | supported | analysis_f20a9f60553ce57f:exec_8eb173511297, analysis_f20a9f60553ce57f:exec_c714034a13f9, analysis_f20a9f60553ce57f:exec_bbf7b5170507, analysis_f20a9f60553ce57f:exec_cf7f251ec694, analysis_f20a9f60553ce57f:exec_25ee09cebebb, analysis_f20a9f60553ce57f:exec_a02b71c30924, analysis_f20a9f60553ce57f:exec_bfffaa01eacb, analysis_f20a9f60553ce57f:exec_6e97406cf140 |
| c23 By static degree, llm/__init__.py is the package's aggregation point (754 incoming calls) and llm/cli.py is a heavy consumer (105 in, 245 out). | inferred | supported | analysis_f20a9f60553ce57f:role_2, analysis_f20a9f60553ce57f:role_4 |
| c24 Within the package, llm/parts.py and llm/models.py carry much of the local call mass, and llm/plugins.py is a leaf: 41 incoming, 0 outgoing. | inferred | supported | analysis_f20a9f60553ce57f:role_12, analysis_f20a9f60553ce57f:role_11, analysis_f20a9f60553ce57f:role_13 |
| c25 The seven dependency records are labeled scope 'runtime' yet all come from docs/requirements.txt (sphinx and documentation tooling), so they describe the docs build toolchain. | observed | supported | analysis_f20a9f60553ce57f:dep_1, analysis_f20a9f60553ce57f:dep_2, analysis_f20a9f60553ce57f:dep_3, analysis_f20a9f60553ce57f:dep_4, analysis_f20a9f60553ce57f:dep_5, analysis_f20a9f60553ce57f:dep_6, analysis_f20a9f60553ce57f:dep_7 |
| c26 lim_3 records that dependency parsing covers only requirements-style manifests and skips pyproject.toml dependency tables, so the package's own runtime dependencies are not established here. | observed | supported | analysis_f20a9f60553ce57f:lim_3, analysis_f20a9f60553ce57f:important_3 |
| c27 Unresolved edges dominate: 7,374 external_or_unresolved versus 2,324 local, and dynamic imports are explicitly not resolved. | inferred | supported | analysis_f20a9f60553ce57f:relation_counts, analysis_f20a9f60553ce57f:lim_1 |
| c28 lim_1 records that reflection, runtime dependency injection, dynamic imports, monkey-patching, generated code, framework runtime wiring, and dynamic dispatch are not resolved by this static analysis. | observed | supported | analysis_f20a9f60553ce57f:lim_1 |
| c29 How plugin loading, model resolution, and logging interact at runtime is not established by this packet. | unresolved | supported | analysis_f20a9f60553ce57f:lim_1, analysis_f20a9f60553ce57f:claim_7c9925f9c8f4 |
| c30 lim_2 records that install/run/build inference is partial and installation commands are not verified. | observed | supported | analysis_f20a9f60553ce57f:lim_2 |
| c31 The packet shows 44 test-like files, but it cannot show whether tests pass or how the code behaves when executed. | inferred | supported | analysis_f20a9f60553ce57f:summary_repository |
| c32 lim_4 records that README claim extraction is line-based and may capture code lines, so README-derived statements stand as author_claimed at most. | observed | supported | analysis_f20a9f60553ce57f:lim_4 |
| c33 A grounded order: README.md and pyproject.toml for framing; llm/__main__.py, llm/cli.py:cli, and llm/plugins.py [load_plugins] for entry and boundary; llm/__init__.py, llm/parts.py, and llm/models.py for the core. | inferred | supported | analysis_f20a9f60553ce57f:important_1, analysis_f20a9f60553ce57f:important_3, analysis_f20a9f60553ce57f:py_entry_8d052da6b4aa, analysis_f20a9f60553ce57f:entry_1, analysis_f20a9f60553ce57f:exec_c714034a13f9, analysis_f20a9f60553ce57f:role_2, analysis_f20a9f60553ce57f:role_12, analysis_f20a9f60553ce57f:role_11 |
| c34 tests/test_parts.py, with 422 outgoing calls and the orchestration label, is a candidate stop for seeing how parts are exercised. | inferred | supported | analysis_f20a9f60553ce57f:role_35 |
| c35 The teaching record recommends starting from manifests and important symbols, then tracing the bounded execution path and inspecting unresolved boundaries before changing code. | inferred | supported | analysis_f20a9f60553ce57f:claim_e56db45165bf |
Sources, rights and disclosure · attribution-license-templates/v0.1
Rights notice. Original repository hosted on GitHub. Repository source code, documentation, names, media, and related project materials remain subject to the rights of their respective authors, contributors, and other rights holders and to applicable repository license terms.
Platform notice. GitHub is the source hosting platform for the linked repository. GitHub and related marks are trademarks of GitHub, Inc. EVEMISS Technology is not affiliated with or endorsed by GitHub unless explicitly stated otherwise.
How this page is produced. This page was produced using revision-aware repository analysis and AI-assisted editorial tooling. Technical claims are tied to the analyzed repository revision and may be revalidated when the source repository changes.
AI-assisted analysis. Reviewed by EVEMISS Technology through human–AI collaborative review. · Report a rights concern · Back to the overview