EVEMISSTechnology

Repository guide · overview · v1

openai / whisper

“Robust Speech Recognition via Large-Scale Weak Supervision” — as described by its authors

  • Artificial Intelligence
  • Python
  • MIT · open-source license
Original repository
openai/whisper
Source platform
GitHub
Repository owner / organization
openai
License
MIT · open-source license
Analyzed revision
86098128c0b4f24f0e2aa2994de830614b474227
Last verified

Source checked 2026-10-03: default branch still at the analyzed revision.

openai/whisper explained: a Python speech-recognition codebase

Terms used on this page
Entrypoint record
A file the analyzer marks as a place where execution can start. When it carries a __main__ guard excerpt, the file contains an if __name__ == "__main__": block and the excerpt shows what that block calls; without an excerpt, the file was flagged by its name only.
Bounded static execution path
A call chain reconstructed from the source code without running it, stopped after a fixed number of steps. It shows how far the code can be followed on paper, not what happens at run time.
Unresolved boundary
Where a static path stops because the next call goes into an external library or cannot be resolved without running the code. It marks the edge of this analysis, not a defect in the repository.
Static relations
Calls and imports found in the source. "External or unresolved" relations point outside the analyzed files.
Module role
The analyzer's label for a file, inferred from how many calls go in and out (for example core, entry/orchestration, leaf). It describes a position in the call graph, not the authors' design.
Test-like files
Files whose names or locations look like tests. This analysis counts them; it does not run them.
Observed · inferred · author-claimed · unresolved
How each statement is supported: read directly from the analyzed files; derived by the analyzer from them; stated by the repository's authors (metadata, README); or not established by this analysis.
Verified
Two uses on these pages. In an analyzer record ("verified provenance", "a verified entrypoint") it means the record was read directly from the analyzed files, which these guides call observed; for an entrypoint, the __main__ guard text is present. Since the analyzer fix of 2026-10-04, a file that only has an entrypoint-like name such as cli.py or main.py is recorded as inferred; guides analyzed before that still call such a file verified, and it is still only a guess about how the file is used. It does not mean the code was run or tested. "Last verified" is the date the guide last passed the lab's checks against the analysis record of the stated revision; it is not a review of the repository itself.

Static overview of the whisper repository: a Python codebase (45 analyzed files) built around a whisper package with model, decoding, timing, tokenizer, audio, and transcribe modules. Its one detected entrypoint is a __main__ guard in whisper/transcribe.py calling cli(). requirements.txt lists seven runtime dependencies and seven test files exist; 1233 of 1388 static relations remain unresolved.

What it is

This is a Python repository: platform metadata records Python as the primary language, with 160526 Python language bytes.

The GitHub description reads "Robust Speech Recognition via Large-Scale Weak Supervision" — the author's own wording, not an analyzed fact.

It is MIT licensed; the latest recorded release is v20250625 (published 2025-06-26), and the metadata snapshot counted 109056 stargazers.

Code sits in a top-level whisper package directory next to a tests directory; the analyzer's summary reports 45 analyzed files, 18 of them under whisper.

How it starts

The one detected script entrypoint is in whisper/transcribe.py:

if __name__ == "__main__":
    cli()

The static core flow is short: cli() (lines 517-619) builds an argparse parser through repeated parser.add_argument calls and checks torch.cuda.is_available.

All six bounded execution paths from __main__ end at these external or unresolved boundaries, so nothing beyond argument setup is traced.

Role analysis labels whisper/transcribe.py an entry/orchestration candidate (3 incoming, 18 outgoing static calls).

Structure

  • Relation totals: 1388 static edges — 1207 calls, 181 imports; 155 resolved locally, 1233 external or unresolved.
  • whisper/model.py: 22 incoming / 20 outgoing — the most connected module, labeled a service/core candidate.
  • whisper/timing.py (16/9), whisper/utils.py (14/2), whisper/tokenizer.py (12/2), whisper/decoding.py (11/13), whisper/audio.py (8/5): also service/core candidates.
  • whisper/__init__.py (3/4) and whisper/normalizers/english.py (3/3) are smaller service/core candidates; tests/test_timing.py (0 in / 6 out) is an orchestration candidate.
  • whisper/transcribe.py (3 in / 18 out) is the entry/orchestration candidate.
  • The pattern suggests the often-called modules (model, timing, utils, tokenizer) form a library core that transcribe.py drives — an inference from call degree, not runtime behavior.

Dependencies and tests

  • Dependency evidence: requirements.txt lists seven runtime dependencies — numba, numpy, torch, tqdm, more-itertools, tiktoken, triton (>=2.0.0, the only version floor).
  • The repository holds two manifests (pyproject.toml, requirements.txt), but only requirements-style files are parsed, so pyproject.toml dependency tables are invisible to this analysis.
  • Tests: 7 test or test-like files, including tests/conftest.py; tests/test_timing.py alone defines test_dtw, test_dtw_cuda_equivalence, test_median_filter, and test_median_filter_equivalence.
  • Install and run guidance here is partial and installation commands are unverified — inspect the manifests before installing.

Read first

  1. README.md — orientation.
  2. pyproject.toml and requirements.txt — intended setup and dependencies.
  3. LICENSE — terms.
  4. whisper/transcribe.py — read transcribe() (lines 38-514), then the cli() parser (517-619).
  5. whisper/__init__.py — load_model(), available_models(), _download().
  6. whisper/model.py — the Whisper class, to connect model construction with transcribe().

The first three steps follow the analyzer's project-level important-file list; whisper/transcribe.py is flagged as a likely entrypoint.

Limits of this analysis

  • Static only: reflection, runtime dependency injection, dynamic imports, monkey-patching, generated code, framework runtime wiring, and dynamic dispatch are not resolved.
  • 1233 of 1388 static relations are external or unresolved.
  • Install/run/build inference is partial, and installation commands were not verified.
  • README extraction is line-based and may capture code lines instead of prose, so README-derived statements are author-claimed, not established.
  • Manifest fields describing the analyzer's own capabilities are excluded from repository evidence.
  • No runtime flow — CLI transcription, model loading, GPU use — is established end to end.

Notable symbols

  • whisper/model.py: Whisper (259-352), AudioEncoder (181-211), TextDecoder (214-256), MultiHeadAttention (81-146), ResidualAttentionBlock (149-178).
  • whisper/decoding.py: detect_language, DecodingOptions, DecodingResult, GreedyDecoder, BeamSearchDecoder.
  • whisper/audio.py: load_audio, pad_or_trim, mel_filters, log_mel_spectrogram; whisper/tokenizer.py: Tokenizer, get_encoding, get_tokenizer.
  • whisper/timing.py: median_filter, dtw_cpu, dtw_cuda, find_alignment, add_word_timestamps.
  • whisper/utils.py: ResultWriter, WriteTXT, SubtitlesWriter, WriteVTT, plus format_timestamp and str2bool.

The AudioEncoder/TextDecoder naming suggests an encoder-decoder split between audio input and text output — a structural inference, not confirmed by execution.

What this analysis could not establish

The writing model's own notes on the analysis record. Identifiers such as lim_3, exec_* or claim_… name records of that analysis; the claims table cites the same records.

  • 1233 of 1388 static relations are external or unresolved, so most cross-module behavior was not traced.
  • pyproject.toml dependency tables were not parsed, so the dependency picture may be incomplete.
  • No code was executed; CLI behavior, model loading, and CUDA paths are unverified.
  • Module roles come from static call degree and entrypoint membership, so they are structural hints, not measured runtime roles.
  • The flagged important file readme.md was excluded because the path is not in the analyzed file inventory (case-variant match).
Claims and evidence — 36 claims, 36 supported by an independent verifier

Every statement above is a claim that cites grounding IDs from the analysis. IDs are internal to the analysis run; the columns show what each claim rests on and how strong that ground is.

Every substantive statement above is a claim bound to grounding IDs of the analyzed revision. Global grounding IDs are namespaced by the analysis run.

Claim Epistemic status Verifier Grounding (global IDs)
c1 Platform metadata records Python as the primary language, with {"Python": 160526} language bytes. observed supported analysis_ee52b965e66bce7b:meta_primary_language, analysis_ee52b965e66bce7b:meta_language_bytes
c2 The repository's GitHub description is "Robust Speech Recognition via Large-Scale Weak Supervision". author_claimed supported analysis_ee52b965e66bce7b:meta_description
c3 The project is MIT licensed, and its latest recorded release is v20250625, published 2025-06-26. observed supported analysis_ee52b965e66bce7b:meta_license, analysis_ee52b965e66bce7b:meta_latest_release, analysis_ee52b965e66bce7b:important_5
c4 Source code sits in a top-level whisper package directory alongside a tests directory. observed supported analysis_ee52b965e66bce7b:subsys_2, analysis_ee52b965e66bce7b:subsys_1
c5 The analyzer's summary reports 45 analyzed files, 18 of them in the whisper directory and 7 in tests. inferred supported analysis_ee52b965e66bce7b:summary_repository
c6 The repository had 109056 stargazers as of 2026-09-14 per platform metadata. observed supported analysis_ee52b965e66bce7b:meta_stars
c7 whisper/transcribe.py contains a __main__ execution guard whose body calls cli(); cli() spans lines 517-619 of that file. observed supported analysis_ee52b965e66bce7b:py_entry_e9d8e13fe4d4, analysis_ee52b965e66bce7b:py_func_a31ce0a9065c
c8 Six bounded static execution paths start at whisper/transcribe.py:__main__ and pass through cli(); each terminates at an external or unresolved call — argparse.ArgumentParser, parser.add_argument (repeatedly), or torch.cuda.is_available. inferred supported analysis_ee52b965e66bce7b:exec_210328c6174a, analysis_ee52b965e66bce7b:exec_873b5bbd14cf, analysis_ee52b965e66bce7b:exec_deca1b3baebb, analysis_ee52b965e66bce7b:exec_210dbd98fcf8, analysis_ee52b965e66bce7b:exec_c8ba6b2b53c8, analysis_ee52b965e66bce7b:exec_bd66f158fc76
c9 The analyzer labels whisper/transcribe.py an entry/orchestration candidate, with 3 incoming and 18 outgoing static calls. inferred supported analysis_ee52b965e66bce7b:role_16
c10 The analyzer recorded 1388 static relations (1207 calls, 181 imports); 155 resolved locally and 1233 stayed external or unresolved. observed supported analysis_ee52b965e66bce7b:relation_counts
c11 whisper/model.py is the most connected module (22 incoming, 20 outgoing calls) and is labeled a service/core candidate. inferred supported analysis_ee52b965e66bce7b:role_10
c12 whisper/decoding.py (11/13), whisper/timing.py (16/9), whisper/tokenizer.py (12/2), whisper/utils.py (14/2), and whisper/audio.py (8/5) are each labeled service/core candidates. inferred supported analysis_ee52b965e66bce7b:role_9, analysis_ee52b965e66bce7b:role_14, analysis_ee52b965e66bce7b:role_15, analysis_ee52b965e66bce7b:role_18, analysis_ee52b965e66bce7b:role_8
c13 whisper/normalizers/english.py (3/3) and whisper/__init__.py (3/4) are smaller service/core candidates; tests/test_timing.py (0 incoming, 6 outgoing) is an orchestration candidate. inferred supported analysis_ee52b965e66bce7b:role_13, analysis_ee52b965e66bce7b:role_6, analysis_ee52b965e66bce7b:role_3
c14 whisper/transcribe.py is labeled an entry/orchestration candidate (3 in / 18 out); combined with the degree numbers, this suggests it drives a core of often-called modules — an inference from call degree, not runtime behavior. inferred supported analysis_ee52b965e66bce7b:role_16, analysis_ee52b965e66bce7b:role_10, analysis_ee52b965e66bce7b:role_14, analysis_ee52b965e66bce7b:role_18, analysis_ee52b965e66bce7b:role_15
c15 requirements.txt lists seven runtime dependencies: numba, numpy, torch, tqdm, more-itertools, tiktoken, and triton (>=2.0.0, the only one with a version floor). observed supported analysis_ee52b965e66bce7b:dep_1, analysis_ee52b965e66bce7b:dep_2, analysis_ee52b965e66bce7b:dep_3, analysis_ee52b965e66bce7b:dep_4, analysis_ee52b965e66bce7b:dep_5, analysis_ee52b965e66bce7b:dep_6, analysis_ee52b965e66bce7b:dep_7
c16 The repository contains 2 build/dependency manifest files, but only requirements-style manifests are parsed; pyproject.toml dependency tables are not. observed supported analysis_ee52b965e66bce7b:ev_manifest_1, analysis_ee52b965e66bce7b:lim_3
c17 The tests directory contains 7 test files or test-like files, including tests/conftest.py. observed supported analysis_ee52b965e66bce7b:ev_tests_1
c18 tests/test_timing.py defines four selected test functions: test_dtw, test_dtw_cuda_equivalence, test_median_filter, and test_median_filter_equivalence. observed supported analysis_ee52b965e66bce7b:py_func_6b8d72c7960f, analysis_ee52b965e66bce7b:py_func_681d1ea53357, analysis_ee52b965e66bce7b:py_func_cf597b3d5d25, analysis_ee52b965e66bce7b:py_func_59835aca4f63
c19 Install guidance from the analyzer is partial: installation commands are not verified. observed supported analysis_ee52b965e66bce7b:lim_2
c20 The packet marks five important files: README.md, pyproject.toml, requirements.txt, LICENSE, and whisper/transcribe.py (flagged as a likely entrypoint). observed supported analysis_ee52b965e66bce7b:important_1, analysis_ee52b965e66bce7b:important_3, analysis_ee52b965e66bce7b:important_4, analysis_ee52b965e66bce7b:important_5, analysis_ee52b965e66bce7b:important_6
c21 A grounded reading order: README.md first for orientation, then pyproject.toml and requirements.txt for setup, then LICENSE, then whisper/transcribe.py. inferred supported analysis_ee52b965e66bce7b:important_1, analysis_ee52b965e66bce7b:important_3, analysis_ee52b965e66bce7b:important_4, analysis_ee52b965e66bce7b:important_5, analysis_ee52b965e66bce7b:important_6
c22 whisper/transcribe.py defines transcribe() at lines 38-514 and cli() at lines 517-619. observed supported analysis_ee52b965e66bce7b:py_func_ed32c16d5c96, analysis_ee52b965e66bce7b:py_func_a31ce0a9065c
c23 whisper/__init__.py defines load_model(), available_models(), and _download(). observed supported analysis_ee52b965e66bce7b:py_func_3caacfc5e3d9, analysis_ee52b965e66bce7b:py_func_7ca1feb87e67, analysis_ee52b965e66bce7b:py_func_a5de7f0ce3d2
c24 After transcribe.py, read the Whisper class in whisper/model.py and load_model() in whisper/__init__.py to connect model construction with the transcription flow. inferred supported analysis_ee52b965e66bce7b:py_class_258632aeaf38, analysis_ee52b965e66bce7b:py_func_3caacfc5e3d9
c25 The analyzer documents that reflection, runtime dependency injection, dynamic imports, monkey-patching, generated code, framework runtime wiring, and dynamic dispatch are not resolved. observed supported analysis_ee52b965e66bce7b:lim_1
c26 1233 of 1388 static relations remain external or unresolved. observed supported analysis_ee52b965e66bce7b:relation_counts
c27 Install/run/build inference is partial and installation commands are not verified. observed supported analysis_ee52b965e66bce7b:lim_2
c28 README claim extraction is line-based and may capture code lines instead of prose, so README-derived statements are author-claimed, not established. observed supported analysis_ee52b965e66bce7b:lim_4
c29 Manifest fields such as project.actual_capabilities describe the analyzer itself, not the analyzed repository, and are excluded from repository evidence. observed supported analysis_ee52b965e66bce7b:lim_5
c30 End-to-end runtime behavior is not established: the packet itself marks dynamic dispatch, reflection, generated-code, and runtime framework behavior as unresolvable statically. unresolved supported analysis_ee52b965e66bce7b:claim_7c9925f9c8f4
c31 whisper/model.py defines Whisper (lines 259-352), AudioEncoder (181-211), TextDecoder (214-256), MultiHeadAttention (81-146), and ResidualAttentionBlock (149-178). observed supported analysis_ee52b965e66bce7b:py_class_258632aeaf38, analysis_ee52b965e66bce7b:py_class_b562f875502f, analysis_ee52b965e66bce7b:py_class_785260aa3f04, analysis_ee52b965e66bce7b:py_class_4ed9b8eada83, analysis_ee52b965e66bce7b:py_class_52b8e7fc7bc9
c32 whisper/decoding.py defines detect_language(), DecodingOptions, DecodingResult, GreedyDecoder, and BeamSearchDecoder, among other classes. observed supported analysis_ee52b965e66bce7b:py_func_aa60011d71a2, analysis_ee52b965e66bce7b:py_class_73648de04931, analysis_ee52b965e66bce7b:py_class_aa1a4e6d4813, analysis_ee52b965e66bce7b:py_class_f015fc72fa5b, analysis_ee52b965e66bce7b:py_class_61d8ff25d720
c33 whisper/audio.py defines load_audio(), pad_or_trim(), mel_filters(), and log_mel_spectrogram(); whisper/tokenizer.py defines Tokenizer, get_encoding(), and get_tokenizer(). observed supported analysis_ee52b965e66bce7b:py_func_0ac35ae1bbef, analysis_ee52b965e66bce7b:py_func_915c49ac6d81, analysis_ee52b965e66bce7b:py_func_499661f28292, analysis_ee52b965e66bce7b:py_func_d200eb51a4b3, analysis_ee52b965e66bce7b:py_class_814237ef1902, analysis_ee52b965e66bce7b:py_func_0d16d45ecbed, analysis_ee52b965e66bce7b:py_func_5be6d8888ea5
c34 whisper/timing.py defines median_filter(), dtw_cpu(), dtw_cuda(), find_alignment(), and add_word_timestamps(). observed supported analysis_ee52b965e66bce7b:py_func_dbcb40294500, analysis_ee52b965e66bce7b:py_func_56b7a739b461, analysis_ee52b965e66bce7b:py_func_fe8c28e8968b, analysis_ee52b965e66bce7b:py_func_e1372f4b817b, analysis_ee52b965e66bce7b:py_func_69d75b442d80
c35 whisper/utils.py defines the writers ResultWriter, WriteTXT, SubtitlesWriter, and WriteVTT, plus helpers such as format_timestamp() and str2bool(). observed supported analysis_ee52b965e66bce7b:py_class_892b197925de, analysis_ee52b965e66bce7b:py_class_16f1ca58bef3, analysis_ee52b965e66bce7b:py_class_e87997a189d5, analysis_ee52b965e66bce7b:py_class_610f01bd8eb9, analysis_ee52b965e66bce7b:py_func_8c9c9c69899c, analysis_ee52b965e66bce7b:py_func_24929df50c11
c36 The AudioEncoder/TextDecoder pair in whisper/model.py suggests an encoder-decoder split between audio input and text output; this is a structural reconstruction, not confirmed by execution. inferred supported analysis_ee52b965e66bce7b:architecture_reconstruction, analysis_ee52b965e66bce7b:py_class_b562f875502f, analysis_ee52b965e66bce7b:py_class_785260aa3f04

Sources, rights and disclosure · attribution-license-templates/v0.1

Rights notice. Original repository hosted on GitHub. Repository source code, documentation, names, media, and related project materials remain subject to the rights of their respective authors, contributors, and other rights holders and to applicable repository license terms.

Platform notice. GitHub is the source hosting platform for the linked repository. GitHub and related marks are trademarks of GitHub, Inc. EVEMISS Technology is not affiliated with or endorsed by GitHub unless explicitly stated otherwise.

How this page is produced. This page was produced using revision-aware repository analysis and AI-assisted editorial tooling. Technical claims are tied to the analyzed repository revision and may be revalidated when the source repository changes.

AI-assisted analysis. Reviewed by EVEMISS Technology through human–AI collaborative review. · Report a rights concern · All repository guides