Repository guide · getting started · v1
simonw / datasette · Getting started
Platform metadata calls it "An open source multi-tool for exploring and publishing data", with Python as the recorded primary language, an Apache-2.0 license and 333 analyzed files. The evidence covers six package.json dependency records, five entrypoint records and 41 bounded execution paths (5 listed), with no verified install or run commands.
- Original repository
- simonw/datasette
- License
- Apache-2.0 · open-source license
- Analyzed revision · last verified
- b338c6f5f6b39dd3e0a341431f071d8cecb0d12b ·
Newer revision observed; parts of this guide may be outdated. The default branch moved to cec5e6b2ef5d, checked 2026-10-03. The whole guide describes revision b338c6f5f6b3. A revision diff found changes in 1 code region it cites (listed below); statements about that region may not hold at the new revision and have not been rechecked yet.
- datasette/app.py:168 · call site on the execution path changed
Getting started with simonw/datasette: metadata and entrypoints
Terms used on this page
- Entrypoint record
- A file the analyzer marks as a place where execution can start. When it carries a __main__ guard excerpt, the file contains an if __name__ == "__main__": block and the excerpt shows what that block calls; without an excerpt, the file was flagged by its name only.
- Bounded static execution path
- A call chain reconstructed from the source code without running it, stopped after a fixed number of steps. It shows how far the code can be followed on paper, not what happens at run time.
- Unresolved boundary
- Where a static path stops because the next call goes into an external library or cannot be resolved without running the code. It marks the edge of this analysis, not a defect in the repository.
- Static relations
- Calls and imports found in the source. "External or unresolved" relations point outside the analyzed files.
- Module role
- The analyzer's label for a file, inferred from how many calls go in and out (for example core, entry/orchestration, leaf). It describes a position in the call graph, not the authors' design.
- Test-like files
- Files whose names or locations look like tests. This analysis counts them; it does not run them.
- Observed · inferred · author-claimed · unresolved
- How each statement is supported: read directly from the analyzed files; derived by the analyzer from them; stated by the repository's authors (metadata, README); or not established by this analysis.
- Verified
- Two uses on these pages. In an analyzer record ("verified provenance", "a verified entrypoint") it means the record was read directly from the analyzed files, which these guides call observed; for an entrypoint, the __main__ guard text is present. Since the analyzer fix of 2026-10-04, a file that only has an entrypoint-like name such as cli.py or main.py is recorded as inferred; guides analyzed before that still call such a file verified, and it is still only a guess about how the file is used. It does not mean the code was run or tested. "Last verified" is the date the guide last passed the lab's checks against the analysis record of the stated revision; it is not a review of the repository itself.
What you are looking at
Platform metadata describes this as "An open source multi-tool for exploring and publishing data", and lists datasette.io as its homepage; both are author-claimed fields, not analyzer findings. The same metadata records the topic tags asgi, automatic-api, csv, datasets, datasette, datasette-io, docker, json, python, sql and sqlite, and names Python the primary language. The license is recorded as Apache-2.0. The latest release record reads "1.0a39 published 2026-09-11T00:05:54Z"; what that label implies is not established here. As size signals, the analyzer's summary counts 333 analyzed files, language-byte totals put Python largest ahead of JavaScript, HTML and CSS, and 11,460 stargazers were recorded on 2026-09-14.
What it needs
One parsed manifest feeds this analysis: package.json — all six dependency records come from it. Five records carry runtime scope: @codemirror/lang-sql ^6.3.3, @rollup/plugin-node-resolve ^15.0.1, @rollup/plugin-terser ^0.1.0, codemirror ^6.0.1 and rollup ^3.30.0. One carries dev scope: prettier ^3.0.0. The analyzer states that dependency records come only from requirements-style manifests parsed by v0.10 and that pyproject.toml dependency tables are not parsed. pyproject.toml is still flagged as a project-level important file, but this analysis does not enumerate what it declares. No installation commands appear here, and the analyzer records install/run/build inference as partial with installation commands not verified.
How it starts
Three entrypoint records are files with a Python __main__ execution guard: datasette/__main__.py, tests/build_small_spatialite_db.py and tests/fixtures.py. The guard in datasette/__main__.py calls cli(); tests/fixtures.py contains a guard calling cli() as well, while tests/build_small_spatialite_db.py calls generate_it("spatialite.db"). Two more records — datasette/app.py and datasette/cli.py — are flagged as likely executable entrypoints by filename heuristic alone. Altogether the packet lists five entrypoint records, and the analyzer reports 41 bounded execution paths, 5 of them listed, covering datasette/__main__.py and datasette/app.py. The listed path from the __main__ guard reaches cli in datasette/cli.py as a local call and stops as a leaf; the four listed datasette/app.py paths stop at unresolved boundaries (Path, logging.getLogger, contextvars.ContextVar, collections.namedtuple). Beyond these records, runtime behaviour is not established: the analysis is static-only and does not resolve dynamic imports or framework wiring.
Where to look first
Project-level important files are README.md, pyproject.toml, package.json and LICENSE. Five further files are flagged as likely entrypoints: datasette/__main__.py, tests/build_small_spatialite_db.py, tests/fixtures.py, datasette/app.py and datasette/cli.py. The analyzer's teaching claims point the way: read README.md early, read datasette/cli.py early, start with manifests and important symbols, trace the bounded execution path and inspect unresolved boundaries before changing code, and begin modification analysis at datasette/app.py, verifying downstream effects manually. Top level, the summary counts datasette (135 files), tests (102), docs (50), .github (17), demos (7) and 22 at the root — 333 in all. Tests exist: the summary reports 102 test-like files, and two guard-bearing scripts noted above sit under tests/.
What this analysis cannot tell you
Commands: install/run/build inference is partial and installation commands are not verified, so nothing above is a run recipe. Dependencies: only requirements-style manifests parsed by v0.10 feed the records, so pyproject.toml's dependency tables are absent. README text: extraction is line-based and may capture code lines instead of prose, so README-derived statements are author-claimed at most. Method: the analysis is static only — reflection, runtime dependency injection, dynamic imports, monkey-patching, generated code, framework runtime wiring and dynamic dispatch are not resolved — and bounded: oversized, minified and vendored files are inventoried but not parsed, with relation extraction capped at 2500 per file and 150000 per repository. Analyzer-capability fields describe the tool, not the repository. Two packet inconsistencies: one teaching claim says no bounded execution path exists while the packet reports 41 (5 listed), and the inferred purpose text says 'web or Node.js application' although the language hint says Python. Runtime behaviour, performance, security and quality characteristics are not established by this analysis.
What the project says about itself
As author-claimed metadata, the repository description reads "An open source multi-tool for exploring and publishing data" and the homepage field lists https://datasette.io; this analysis neither verifies nor extends those fields.
What this analysis could not establish
The writing model's own notes on the analysis record. Identifiers such as lim_3, exec_* or claim_… name records of that analysis; the claims table cites the same records.
- Only 5 of the 41 reported bounded execution paths are listed, so the startup picture beyond datasette/__main__.py and datasette/app.py is incomplete.
- pyproject.toml dependency tables are not parsed (lim_3), so Python-side dependencies are unknown to this analysis.
- Teaching claim claim_4d30ade3e21e says no bounded static execution path is available, contradicting the 41-path total reported elsewhere in the same packet.
- The analyzer's inferred purpose ('web or Node.js application') conflicts with its own Python-primary language hint and was treated as a generic guess.
- No install, build or run commands are verified by this analysis, so the asset deliberately contains none.
Claims and evidence — 35 claims, 35 supported by an independent verifier
Every substantive statement above is a claim bound to grounding IDs of the analyzed revision. Global grounding IDs are namespaced by the analysis run.
| Claim | Epistemic status | Verifier | Grounding (global IDs) |
|---|---|---|---|
| c1 Platform metadata describes the project as "An open source multi-tool for exploring and publishing data" and lists datasette.io as its homepage; both are author-claimed fields rather than analyzer findings. | author_claimed | supported | analysis_ef14c7222eb9f538:meta_description, analysis_ef14c7222eb9f538:meta_homepage |
| c2 The recorded topic tags are asgi, automatic-api, csv, datasets, datasette, datasette-io, docker, json, python, sql and sqlite, and Python is recorded as the primary language. | observed | supported | analysis_ef14c7222eb9f538:meta_topics, analysis_ef14c7222eb9f538:meta_primary_language |
| c3 The platform records the license as Apache-2.0. | observed | supported | analysis_ef14c7222eb9f538:meta_license |
| c4 The latest release record reads "1.0a39 published 2026-09-11T00:05:54Z"; the packet does not establish what that label implies. | observed | supported | analysis_ef14c7222eb9f538:meta_latest_release |
| c5 The analyzer's summary counts 333 analyzed files; language-byte counts put Python largest (about 2.47 million bytes) ahead of JavaScript and HTML; 11,460 stargazers were recorded at 2026-09-14. | inferred | supported | analysis_ef14c7222eb9f538:summary_repository, analysis_ef14c7222eb9f538:meta_language_bytes, analysis_ef14c7222eb9f538:meta_stars |
| c7 The only parsed dependency manifest is package.json, and all six dependency records come from it. | observed | supported | analysis_ef14c7222eb9f538:dep_1, analysis_ef14c7222eb9f538:dep_2, analysis_ef14c7222eb9f538:dep_3, analysis_ef14c7222eb9f538:dep_4, analysis_ef14c7222eb9f538:dep_5, analysis_ef14c7222eb9f538:dep_6 |
| c8 Five dependency records carry runtime scope: @codemirror/lang-sql ^6.3.3, @rollup/plugin-node-resolve ^15.0.1, @rollup/plugin-terser ^0.1.0, codemirror ^6.0.1 and rollup ^3.30.0. | observed | supported | analysis_ef14c7222eb9f538:dep_1, analysis_ef14c7222eb9f538:dep_2, analysis_ef14c7222eb9f538:dep_3, analysis_ef14c7222eb9f538:dep_4, analysis_ef14c7222eb9f538:dep_5 |
| c9 One dependency record carries dev scope: prettier ^3.0.0. | observed | supported | analysis_ef14c7222eb9f538:dep_6 |
| c10 The analyzer's limitation record states dependency records come only from requirements-style manifests parsed by v0.10 and that pyproject.toml dependency tables are not parsed. | observed | supported | analysis_ef14c7222eb9f538:lim_3 |
| c11 pyproject.toml is flagged as a project-level important file, but this analysis does not enumerate the dependency tables it contains. | observed | supported | analysis_ef14c7222eb9f538:important_3, analysis_ef14c7222eb9f538:lim_3 |
| c12 The analyzer records install/run/build inference as partial and installation commands as not verified; no installation commands appear in this asset. | observed | supported | analysis_ef14c7222eb9f538:lim_2 |
| c13 Three entrypoint records describe files with a Python __main__ execution guard: datasette/__main__.py, tests/build_small_spatialite_db.py and tests/fixtures.py. | observed | supported | analysis_ef14c7222eb9f538:py_entry_7cbf1557a33b, analysis_ef14c7222eb9f538:py_entry_f23c245ed3b2, analysis_ef14c7222eb9f538:py_entry_d7853af71ceb |
c14 The datasette/__main__.py guard calls cli(); the tests/fixtures.py guard calls cli() as well; the tests/build_small_spatialite_db.py guard calls generate_it("spatialite.db"). |
observed | supported | analysis_ef14c7222eb9f538:py_entry_7cbf1557a33b, analysis_ef14c7222eb9f538:py_entry_d7853af71ceb, analysis_ef14c7222eb9f538:py_entry_f23c245ed3b2 |
| c15 datasette/app.py and datasette/cli.py are flagged as likely executable entrypoints by filename heuristic, with no guard excerpt in their records. | observed | supported | analysis_ef14c7222eb9f538:entry_1, analysis_ef14c7222eb9f538:entry_2 |
| c16 The packet lists five entrypoint records in total, and the analyzer reports 41 bounded execution paths, of which 5 are listed, covering datasette/__main__.py and datasette/app.py. | inferred | supported | analysis_ef14c7222eb9f538:summary_repository, analysis_ef14c7222eb9f538:claim_036e282c2657 |
| c17 The listed path from the datasette/__main__.py guard reaches cli in datasette/cli.py as a local call and stops as a leaf; four listed datasette/app.py paths stop at unresolved boundaries: Path, logging.getLogger, contextvars.ContextVar and collections.namedtuple. | inferred | supported | analysis_ef14c7222eb9f538:exec_707222db7901, analysis_ef14c7222eb9f538:exec_bb193f5bffdb, analysis_ef14c7222eb9f538:exec_141532159c69, analysis_ef14c7222eb9f538:exec_6c1e96f0a6c1, analysis_ef14c7222eb9f538:exec_99dfaf114e63 |
| c18 Runtime behaviour beyond these entrypoint and path records is not established: the analysis is static-only and does not resolve dynamic imports, framework runtime wiring or dynamic dispatch. | observed | supported | analysis_ef14c7222eb9f538:lim_1 |
| c19 The analyzer's learning-path claim: start with manifests and important symbols, then trace the bounded execution path and inspect unresolved boundaries before changing code. | inferred | supported | analysis_ef14c7222eb9f538:claim_e56db45165bf |
| c20 The analyzer's modify-guide claim: begin modification analysis at the detected entrypoint datasette/app.py, then verify downstream effects manually. | inferred | supported | analysis_ef14c7222eb9f538:claim_bbaf911cc786 |
| c21 Project-level important files are README.md, pyproject.toml, package.json and LICENSE. | observed | supported | analysis_ef14c7222eb9f538:important_1, analysis_ef14c7222eb9f538:important_3, analysis_ef14c7222eb9f538:important_4, analysis_ef14c7222eb9f538:important_5 |
| c22 Files flagged as important for being likely entrypoints are datasette/__main__.py, tests/build_small_spatialite_db.py, tests/fixtures.py, datasette/app.py and datasette/cli.py. | observed | supported | analysis_ef14c7222eb9f538:important_6, analysis_ef14c7222eb9f538:important_7, analysis_ef14c7222eb9f538:important_8, analysis_ef14c7222eb9f538:important_9, analysis_ef14c7222eb9f538:important_10 |
| c23 Teaching claims recommend reading README.md early as a project-level important file and reading datasette/cli.py early as a likely entrypoint. | inferred | supported | analysis_ef14c7222eb9f538:claim_5e040738cac5, analysis_ef14c7222eb9f538:claim_3cd5a6c42c2e |
| c24 The summary counts 333 files across top-level directories: datasette (135), tests (102), docs (50), .github (17), demos (7) and 22 at the root. | inferred | supported | analysis_ef14c7222eb9f538:summary_repository |
| c25 The summary reports 102 test-like files, so tests exist; tests/build_small_spatialite_db.py and tests/fixtures.py are test-directory files that also carry __main__ guards. | inferred | supported | analysis_ef14c7222eb9f538:summary_repository, analysis_ef14c7222eb9f538:py_entry_f23c245ed3b2, analysis_ef14c7222eb9f538:py_entry_d7853af71ceb |
| c26 The analyzer records install/run/build inference as partial; installation commands are not verified, so no install or run commands are established by this analysis. | observed | supported | analysis_ef14c7222eb9f538:lim_2 |
| c27 Only requirements-style manifests parsed by v0.10 feed dependency records, so pyproject.toml dependency tables are absent from this analysis. | observed | supported | analysis_ef14c7222eb9f538:lim_3 |
| c28 README claim extraction is line-based and may capture code lines instead of prose; README-derived statements are author-claimed at most. | observed | supported | analysis_ef14c7222eb9f538:lim_4 |
| c29 The analysis is static only: reflection, runtime dependency injection, dynamic imports, monkey-patching, generated code, framework runtime wiring and dynamic dispatch are not resolved. | observed | supported | analysis_ef14c7222eb9f538:lim_1 |
| c30 Analysis is bounded: files over 400 kB, minified bundles and vendored directories are inventoried but not parsed; relation extraction is capped at 2500 per file and 150000 per repository. | observed | supported | analysis_ef14c7222eb9f538:lim_6 |
| c36 The packet's analyzer-capability fields describe the analyzer, not the analyzed repository, and are excluded from repository evidence. | observed | supported | analysis_ef14c7222eb9f538:lim_5 |
| c31 Packet inconsistency: one teaching claim states no bounded static execution path is available and runtime flow is unresolved, while the same packet reports 41 bounded execution paths with 5 listed; this asset treats the path records as operative. | inferred | supported | analysis_ef14c7222eb9f538:claim_4d30ade3e21e, analysis_ef14c7222eb9f538:summary_repository |
| c32 Packet inconsistency: the analyzer's inferred purpose text calls the repository a 'web or Node.js application' while its language hint and the platform's primary-language record say Python; treat the purpose line as a generic guess. | inferred | supported | analysis_ef14c7222eb9f538:summary_repository, analysis_ef14c7222eb9f538:meta_primary_language |
| c33 Runtime behaviour, performance, security and quality characteristics of the project are not established by this static analysis. | unresolved | supported | analysis_ef14c7222eb9f538:lim_1 |
| c34 The author-claimed repository description reads "An open source multi-tool for exploring and publishing data". | author_claimed | supported | analysis_ef14c7222eb9f538:meta_description |
| c35 The author-claimed homepage field lists https://datasette.io. | author_claimed | supported | analysis_ef14c7222eb9f538:meta_homepage |
Sources, rights and disclosure · attribution-license-templates/v0.1
Rights notice. Original repository hosted on GitHub. Repository source code, documentation, names, media, and related project materials remain subject to the rights of their respective authors, contributors, and other rights holders and to applicable repository license terms.
Platform notice. GitHub is the source hosting platform for the linked repository. GitHub and related marks are trademarks of GitHub, Inc. EVEMISS Technology is not affiliated with or endorsed by GitHub unless explicitly stated otherwise.
How this page is produced. This page was produced using revision-aware repository analysis and AI-assisted editorial tooling. Technical claims are tied to the analyzed repository revision and may be revalidated when the source repository changes.
AI-assisted analysis. Reviewed by EVEMISS Technology through human–AI collaborative review. · Report a rights concern · Back to the overview