Repository guide · overview · v1
openai / whisper
“Robust Speech Recognition via Large-Scale Weak Supervision” — as described by its authors
- Artificial Intelligence
- Python
- MIT · open-source license
- Original repository
- openai/whisper
- Source platform
- GitHub
- Repository owner / organization
- openai
- License
- MIT · open-source license
- Analyzed revision
- 86098128c0b4f24f0e2aa2994de830614b474227
- Last verified
Source checked 2026-10-03: default branch still at the analyzed revision.
openai/whisper explained: a Python speech-recognition codebase
Terms used on this page
- Entrypoint record
- A file the analyzer marks as a place where execution can start. When it carries a __main__ guard excerpt, the file contains an if __name__ == "__main__": block and the excerpt shows what that block calls; without an excerpt, the file was flagged by its name only.
- Bounded static execution path
- A call chain reconstructed from the source code without running it, stopped after a fixed number of steps. It shows how far the code can be followed on paper, not what happens at run time.
- Unresolved boundary
- Where a static path stops because the next call goes into an external library or cannot be resolved without running the code. It marks the edge of this analysis, not a defect in the repository.
- Static relations
- Calls and imports found in the source. "External or unresolved" relations point outside the analyzed files.
- Module role
- The analyzer's label for a file, inferred from how many calls go in and out (for example core, entry/orchestration, leaf). It describes a position in the call graph, not the authors' design.
- Test-like files
- Files whose names or locations look like tests. This analysis counts them; it does not run them.
- Observed · inferred · author-claimed · unresolved
- How each statement is supported: read directly from the analyzed files; derived by the analyzer from them; stated by the repository's authors (metadata, README); or not established by this analysis.
- Verified
- Two uses on these pages. In an analyzer record ("verified provenance", "a verified entrypoint") it means the record was read directly from the analyzed files, which these guides call observed; for an entrypoint, the __main__ guard text is present. Since the analyzer fix of 2026-10-04, a file that only has an entrypoint-like name such as cli.py or main.py is recorded as inferred; guides analyzed before that still call such a file verified, and it is still only a guess about how the file is used. It does not mean the code was run or tested. "Last verified" is the date the guide last passed the lab's checks against the analysis record of the stated revision; it is not a review of the repository itself.
Static overview of the whisper repository: a Python codebase (45 analyzed files) built around a whisper package with model, decoding, timing, tokenizer, audio, and transcribe modules. Its one detected entrypoint is a __main__ guard in whisper/transcribe.py calling cli(). requirements.txt lists seven runtime dependencies and seven test files exist; 1233 of 1388 static relations remain unresolved.
What it is
This is a Python repository: platform metadata records Python as the primary language, with 160526 Python language bytes.
The GitHub description reads "Robust Speech Recognition via Large-Scale Weak Supervision" — the author's own wording, not an analyzed fact.
It is MIT licensed; the latest recorded release is v20250625 (published 2025-06-26), and the metadata snapshot counted 109056 stargazers.
Code sits in a top-level whisper package directory next to a tests directory; the analyzer's summary reports 45 analyzed files, 18 of them under whisper.
How it starts
The one detected script entrypoint is in whisper/transcribe.py:
if __name__ == "__main__":
cli()
The static core flow is short: cli() (lines 517-619) builds an argparse parser through repeated parser.add_argument calls and checks torch.cuda.is_available.
All six bounded execution paths from __main__ end at these external or unresolved boundaries, so nothing beyond argument setup is traced.
Role analysis labels whisper/transcribe.py an entry/orchestration candidate (3 incoming, 18 outgoing static calls).
Structure
- Relation totals: 1388 static edges — 1207 calls, 181 imports; 155 resolved locally, 1233 external or unresolved.
whisper/model.py: 22 incoming / 20 outgoing — the most connected module, labeled a service/core candidate.whisper/timing.py(16/9),whisper/utils.py(14/2),whisper/tokenizer.py(12/2),whisper/decoding.py(11/13),whisper/audio.py(8/5): also service/core candidates.whisper/__init__.py(3/4) andwhisper/normalizers/english.py(3/3) are smaller service/core candidates;tests/test_timing.py(0 in / 6 out) is an orchestration candidate.whisper/transcribe.py(3 in / 18 out) is the entry/orchestration candidate.- The pattern suggests the often-called modules (
model,timing,utils,tokenizer) form a library core thattranscribe.pydrives — an inference from call degree, not runtime behavior.
Dependencies and tests
- Dependency evidence:
requirements.txtlists seven runtime dependencies — numba, numpy, torch, tqdm, more-itertools, tiktoken, triton (>=2.0.0, the only version floor). - The repository holds two manifests (
pyproject.toml,requirements.txt), but only requirements-style files are parsed, sopyproject.tomldependency tables are invisible to this analysis. - Tests: 7 test or test-like files, including
tests/conftest.py;tests/test_timing.pyalone definestest_dtw,test_dtw_cuda_equivalence,test_median_filter, andtest_median_filter_equivalence. - Install and run guidance here is partial and installation commands are unverified — inspect the manifests before installing.
Read first
README.md— orientation.pyproject.tomlandrequirements.txt— intended setup and dependencies.LICENSE— terms.whisper/transcribe.py— readtranscribe()(lines 38-514), then thecli()parser (517-619).whisper/__init__.py—load_model(),available_models(),_download().whisper/model.py— theWhisperclass, to connect model construction withtranscribe().
The first three steps follow the analyzer's project-level important-file list; whisper/transcribe.py is flagged as a likely entrypoint.
Limits of this analysis
- Static only: reflection, runtime dependency injection, dynamic imports, monkey-patching, generated code, framework runtime wiring, and dynamic dispatch are not resolved.
- 1233 of 1388 static relations are external or unresolved.
- Install/run/build inference is partial, and installation commands were not verified.
- README extraction is line-based and may capture code lines instead of prose, so README-derived statements are author-claimed, not established.
- Manifest fields describing the analyzer's own capabilities are excluded from repository evidence.
- No runtime flow — CLI transcription, model loading, GPU use — is established end to end.
Notable symbols
whisper/model.py:Whisper(259-352),AudioEncoder(181-211),TextDecoder(214-256),MultiHeadAttention(81-146),ResidualAttentionBlock(149-178).whisper/decoding.py:detect_language,DecodingOptions,DecodingResult,GreedyDecoder,BeamSearchDecoder.whisper/audio.py:load_audio,pad_or_trim,mel_filters,log_mel_spectrogram;whisper/tokenizer.py:Tokenizer,get_encoding,get_tokenizer.whisper/timing.py:median_filter,dtw_cpu,dtw_cuda,find_alignment,add_word_timestamps.whisper/utils.py:ResultWriter,WriteTXT,SubtitlesWriter,WriteVTT, plusformat_timestampandstr2bool.
The AudioEncoder/TextDecoder naming suggests an encoder-decoder split between audio input and text output — a structural inference, not confirmed by execution.
What this analysis could not establish
The writing model's own notes on the analysis record. Identifiers such as lim_3, exec_* or claim_… name records of that analysis; the claims table cites the same records.
- 1233 of 1388 static relations are external or unresolved, so most cross-module behavior was not traced.
- pyproject.toml dependency tables were not parsed, so the dependency picture may be incomplete.
- No code was executed; CLI behavior, model loading, and CUDA paths are unverified.
- Module roles come from static call degree and entrypoint membership, so they are structural hints, not measured runtime roles.
- The flagged important file readme.md was excluded because the path is not in the analyzed file inventory (case-variant match).
Claims and evidence — 36 claims, 36 supported by an independent verifier
Every statement above is a claim that cites grounding IDs from the analysis. IDs are internal to the analysis run; the columns show what each claim rests on and how strong that ground is.
Every substantive statement above is a claim bound to grounding IDs of the analyzed revision. Global grounding IDs are namespaced by the analysis run.
| Claim | Epistemic status | Verifier | Grounding (global IDs) |
|---|---|---|---|
| c1 Platform metadata records Python as the primary language, with {"Python": 160526} language bytes. | observed | supported | analysis_ee52b965e66bce7b:meta_primary_language, analysis_ee52b965e66bce7b:meta_language_bytes |
| c2 The repository's GitHub description is "Robust Speech Recognition via Large-Scale Weak Supervision". | author_claimed | supported | analysis_ee52b965e66bce7b:meta_description |
| c3 The project is MIT licensed, and its latest recorded release is v20250625, published 2025-06-26. | observed | supported | analysis_ee52b965e66bce7b:meta_license, analysis_ee52b965e66bce7b:meta_latest_release, analysis_ee52b965e66bce7b:important_5 |
| c4 Source code sits in a top-level whisper package directory alongside a tests directory. | observed | supported | analysis_ee52b965e66bce7b:subsys_2, analysis_ee52b965e66bce7b:subsys_1 |
| c5 The analyzer's summary reports 45 analyzed files, 18 of them in the whisper directory and 7 in tests. | inferred | supported | analysis_ee52b965e66bce7b:summary_repository |
| c6 The repository had 109056 stargazers as of 2026-09-14 per platform metadata. | observed | supported | analysis_ee52b965e66bce7b:meta_stars |
| c7 whisper/transcribe.py contains a __main__ execution guard whose body calls cli(); cli() spans lines 517-619 of that file. | observed | supported | analysis_ee52b965e66bce7b:py_entry_e9d8e13fe4d4, analysis_ee52b965e66bce7b:py_func_a31ce0a9065c |
| c8 Six bounded static execution paths start at whisper/transcribe.py:__main__ and pass through cli(); each terminates at an external or unresolved call — argparse.ArgumentParser, parser.add_argument (repeatedly), or torch.cuda.is_available. | inferred | supported | analysis_ee52b965e66bce7b:exec_210328c6174a, analysis_ee52b965e66bce7b:exec_873b5bbd14cf, analysis_ee52b965e66bce7b:exec_deca1b3baebb, analysis_ee52b965e66bce7b:exec_210dbd98fcf8, analysis_ee52b965e66bce7b:exec_c8ba6b2b53c8, analysis_ee52b965e66bce7b:exec_bd66f158fc76 |
| c9 The analyzer labels whisper/transcribe.py an entry/orchestration candidate, with 3 incoming and 18 outgoing static calls. | inferred | supported | analysis_ee52b965e66bce7b:role_16 |
| c10 The analyzer recorded 1388 static relations (1207 calls, 181 imports); 155 resolved locally and 1233 stayed external or unresolved. | observed | supported | analysis_ee52b965e66bce7b:relation_counts |
| c11 whisper/model.py is the most connected module (22 incoming, 20 outgoing calls) and is labeled a service/core candidate. | inferred | supported | analysis_ee52b965e66bce7b:role_10 |
| c12 whisper/decoding.py (11/13), whisper/timing.py (16/9), whisper/tokenizer.py (12/2), whisper/utils.py (14/2), and whisper/audio.py (8/5) are each labeled service/core candidates. | inferred | supported | analysis_ee52b965e66bce7b:role_9, analysis_ee52b965e66bce7b:role_14, analysis_ee52b965e66bce7b:role_15, analysis_ee52b965e66bce7b:role_18, analysis_ee52b965e66bce7b:role_8 |
| c13 whisper/normalizers/english.py (3/3) and whisper/__init__.py (3/4) are smaller service/core candidates; tests/test_timing.py (0 incoming, 6 outgoing) is an orchestration candidate. | inferred | supported | analysis_ee52b965e66bce7b:role_13, analysis_ee52b965e66bce7b:role_6, analysis_ee52b965e66bce7b:role_3 |
| c14 whisper/transcribe.py is labeled an entry/orchestration candidate (3 in / 18 out); combined with the degree numbers, this suggests it drives a core of often-called modules — an inference from call degree, not runtime behavior. | inferred | supported | analysis_ee52b965e66bce7b:role_16, analysis_ee52b965e66bce7b:role_10, analysis_ee52b965e66bce7b:role_14, analysis_ee52b965e66bce7b:role_18, analysis_ee52b965e66bce7b:role_15 |
| c15 requirements.txt lists seven runtime dependencies: numba, numpy, torch, tqdm, more-itertools, tiktoken, and triton (>=2.0.0, the only one with a version floor). | observed | supported | analysis_ee52b965e66bce7b:dep_1, analysis_ee52b965e66bce7b:dep_2, analysis_ee52b965e66bce7b:dep_3, analysis_ee52b965e66bce7b:dep_4, analysis_ee52b965e66bce7b:dep_5, analysis_ee52b965e66bce7b:dep_6, analysis_ee52b965e66bce7b:dep_7 |
| c16 The repository contains 2 build/dependency manifest files, but only requirements-style manifests are parsed; pyproject.toml dependency tables are not. | observed | supported | analysis_ee52b965e66bce7b:ev_manifest_1, analysis_ee52b965e66bce7b:lim_3 |
| c17 The tests directory contains 7 test files or test-like files, including tests/conftest.py. | observed | supported | analysis_ee52b965e66bce7b:ev_tests_1 |
| c18 tests/test_timing.py defines four selected test functions: test_dtw, test_dtw_cuda_equivalence, test_median_filter, and test_median_filter_equivalence. | observed | supported | analysis_ee52b965e66bce7b:py_func_6b8d72c7960f, analysis_ee52b965e66bce7b:py_func_681d1ea53357, analysis_ee52b965e66bce7b:py_func_cf597b3d5d25, analysis_ee52b965e66bce7b:py_func_59835aca4f63 |
| c19 Install guidance from the analyzer is partial: installation commands are not verified. | observed | supported | analysis_ee52b965e66bce7b:lim_2 |
| c20 The packet marks five important files: README.md, pyproject.toml, requirements.txt, LICENSE, and whisper/transcribe.py (flagged as a likely entrypoint). | observed | supported | analysis_ee52b965e66bce7b:important_1, analysis_ee52b965e66bce7b:important_3, analysis_ee52b965e66bce7b:important_4, analysis_ee52b965e66bce7b:important_5, analysis_ee52b965e66bce7b:important_6 |
| c21 A grounded reading order: README.md first for orientation, then pyproject.toml and requirements.txt for setup, then LICENSE, then whisper/transcribe.py. | inferred | supported | analysis_ee52b965e66bce7b:important_1, analysis_ee52b965e66bce7b:important_3, analysis_ee52b965e66bce7b:important_4, analysis_ee52b965e66bce7b:important_5, analysis_ee52b965e66bce7b:important_6 |
| c22 whisper/transcribe.py defines transcribe() at lines 38-514 and cli() at lines 517-619. | observed | supported | analysis_ee52b965e66bce7b:py_func_ed32c16d5c96, analysis_ee52b965e66bce7b:py_func_a31ce0a9065c |
| c23 whisper/__init__.py defines load_model(), available_models(), and _download(). | observed | supported | analysis_ee52b965e66bce7b:py_func_3caacfc5e3d9, analysis_ee52b965e66bce7b:py_func_7ca1feb87e67, analysis_ee52b965e66bce7b:py_func_a5de7f0ce3d2 |
| c24 After transcribe.py, read the Whisper class in whisper/model.py and load_model() in whisper/__init__.py to connect model construction with the transcription flow. | inferred | supported | analysis_ee52b965e66bce7b:py_class_258632aeaf38, analysis_ee52b965e66bce7b:py_func_3caacfc5e3d9 |
| c25 The analyzer documents that reflection, runtime dependency injection, dynamic imports, monkey-patching, generated code, framework runtime wiring, and dynamic dispatch are not resolved. | observed | supported | analysis_ee52b965e66bce7b:lim_1 |
| c26 1233 of 1388 static relations remain external or unresolved. | observed | supported | analysis_ee52b965e66bce7b:relation_counts |
| c27 Install/run/build inference is partial and installation commands are not verified. | observed | supported | analysis_ee52b965e66bce7b:lim_2 |
| c28 README claim extraction is line-based and may capture code lines instead of prose, so README-derived statements are author-claimed, not established. | observed | supported | analysis_ee52b965e66bce7b:lim_4 |
| c29 Manifest fields such as project.actual_capabilities describe the analyzer itself, not the analyzed repository, and are excluded from repository evidence. | observed | supported | analysis_ee52b965e66bce7b:lim_5 |
| c30 End-to-end runtime behavior is not established: the packet itself marks dynamic dispatch, reflection, generated-code, and runtime framework behavior as unresolvable statically. | unresolved | supported | analysis_ee52b965e66bce7b:claim_7c9925f9c8f4 |
| c31 whisper/model.py defines Whisper (lines 259-352), AudioEncoder (181-211), TextDecoder (214-256), MultiHeadAttention (81-146), and ResidualAttentionBlock (149-178). | observed | supported | analysis_ee52b965e66bce7b:py_class_258632aeaf38, analysis_ee52b965e66bce7b:py_class_b562f875502f, analysis_ee52b965e66bce7b:py_class_785260aa3f04, analysis_ee52b965e66bce7b:py_class_4ed9b8eada83, analysis_ee52b965e66bce7b:py_class_52b8e7fc7bc9 |
| c32 whisper/decoding.py defines detect_language(), DecodingOptions, DecodingResult, GreedyDecoder, and BeamSearchDecoder, among other classes. | observed | supported | analysis_ee52b965e66bce7b:py_func_aa60011d71a2, analysis_ee52b965e66bce7b:py_class_73648de04931, analysis_ee52b965e66bce7b:py_class_aa1a4e6d4813, analysis_ee52b965e66bce7b:py_class_f015fc72fa5b, analysis_ee52b965e66bce7b:py_class_61d8ff25d720 |
| c33 whisper/audio.py defines load_audio(), pad_or_trim(), mel_filters(), and log_mel_spectrogram(); whisper/tokenizer.py defines Tokenizer, get_encoding(), and get_tokenizer(). | observed | supported | analysis_ee52b965e66bce7b:py_func_0ac35ae1bbef, analysis_ee52b965e66bce7b:py_func_915c49ac6d81, analysis_ee52b965e66bce7b:py_func_499661f28292, analysis_ee52b965e66bce7b:py_func_d200eb51a4b3, analysis_ee52b965e66bce7b:py_class_814237ef1902, analysis_ee52b965e66bce7b:py_func_0d16d45ecbed, analysis_ee52b965e66bce7b:py_func_5be6d8888ea5 |
| c34 whisper/timing.py defines median_filter(), dtw_cpu(), dtw_cuda(), find_alignment(), and add_word_timestamps(). | observed | supported | analysis_ee52b965e66bce7b:py_func_dbcb40294500, analysis_ee52b965e66bce7b:py_func_56b7a739b461, analysis_ee52b965e66bce7b:py_func_fe8c28e8968b, analysis_ee52b965e66bce7b:py_func_e1372f4b817b, analysis_ee52b965e66bce7b:py_func_69d75b442d80 |
| c35 whisper/utils.py defines the writers ResultWriter, WriteTXT, SubtitlesWriter, and WriteVTT, plus helpers such as format_timestamp() and str2bool(). | observed | supported | analysis_ee52b965e66bce7b:py_class_892b197925de, analysis_ee52b965e66bce7b:py_class_16f1ca58bef3, analysis_ee52b965e66bce7b:py_class_e87997a189d5, analysis_ee52b965e66bce7b:py_class_610f01bd8eb9, analysis_ee52b965e66bce7b:py_func_8c9c9c69899c, analysis_ee52b965e66bce7b:py_func_24929df50c11 |
| c36 The AudioEncoder/TextDecoder pair in whisper/model.py suggests an encoder-decoder split between audio input and text output; this is a structural reconstruction, not confirmed by execution. | inferred | supported | analysis_ee52b965e66bce7b:architecture_reconstruction, analysis_ee52b965e66bce7b:py_class_b562f875502f, analysis_ee52b965e66bce7b:py_class_785260aa3f04 |
Sources, rights and disclosure · attribution-license-templates/v0.1
Rights notice. Original repository hosted on GitHub. Repository source code, documentation, names, media, and related project materials remain subject to the rights of their respective authors, contributors, and other rights holders and to applicable repository license terms.
Platform notice. GitHub is the source hosting platform for the linked repository. GitHub and related marks are trademarks of GitHub, Inc. EVEMISS Technology is not affiliated with or endorsed by GitHub unless explicitly stated otherwise.
How this page is produced. This page was produced using revision-aware repository analysis and AI-assisted editorial tooling. Technical claims are tied to the analyzed repository revision and may be revalidated when the source repository changes.
AI-assisted analysis. Reviewed by EVEMISS Technology through human–AI collaborative review. · Report a rights concern · All repository guides