EVEMISSTechnology

Repository 指南 · 概覽 · v1

openai / whisper

“Robust Speech Recognition via Large-Scale Weak Supervision” — 作者自述

  • Artificial Intelligence
  • Python
  • MIT · 開源授權
原始 repository
openai/whisper
來源平台
GitHub
擁有者 / 組織
openai
授權
MIT · 開源授權
分析的版本
86098128c0b4f24f0e2aa2994de830614b474227
最後驗證

2026-10-03 檢查過原始專案:預設分支仍在分析的版本。

本頁中文由負責的 AI 編輯依英文正式版本 v1 翻譯;程式碼名稱、路徑、行號與數字都經確定性檢查,與英文版一致。陳述與依據表保留英文原文,因為那是獨立驗證者核對過的紀錄。看英文原文

openai/whisper 解說:一個 Python 語音辨識程式碼庫

本頁用語說明
Entrypoint record(進入點記錄)
分析器標記為「程式可能從這裡開始執行」的檔案。附有 __main__ guard 摘錄時,表示檔案裡有 if __name__ == "__main__": 區塊,摘錄顯示它呼叫了什麼;沒有摘錄時,只是依檔名判斷。
Bounded static execution path(有界靜態執行路徑)
不執行程式、直接從原始碼重建出的呼叫鏈,走到固定步數就停。它說明在紙面上能追到多遠,不代表實際執行時會發生什麼。
Unresolved boundary(未解析邊界)
靜態路徑停下的地方:下一個呼叫進入外部函式庫,或不執行就無法確定。這是本次分析的邊界,不是 repository 的缺陷。
Static relations(靜態關係)
在原始碼中找到的呼叫與 import。「external or unresolved」表示指向分析範圍以外。
Module role(模組角色)
分析器依呼叫進出次數推論的檔案標籤(例如 core、entry/orchestration、leaf)。它描述的是在呼叫圖中的位置,不是作者的設計意圖。
Test-like files(類測試檔案)
名稱或位置看起來像測試的檔案。本次分析只計數,不執行。
Observed · inferred · author-claimed · unresolved
每句話的依據:直接讀自分析的檔案;由分析器從檔案推論;repository 作者自述(metadata、README);或本次分析無法確定。
Verified(已驗證)
本頁有兩種用法。在分析器的記錄裡(「verified provenance」「verified entrypoint」),它表示這筆記錄是直接從分析的檔案讀到的,也就是本指南所說的 observed;以進入點來說,就是檔案裡確實有 __main__ guard 的文字。分析器在 2026-10-04 修正後,只有 cli.py、main.py 這類像進入點的檔名、沒有 guard 的檔案,會記為 inferred(推測);在那之前分析的指南仍把這類檔案標成 verified,它依然只是對檔案用途的推測。它不代表程式被執行或測試過。「最後驗證」是這份指南最後一次通過本實驗室檢查的日期,檢查對象是所標示版本的分析記錄,不是對 repository 本身的審查。

whisper repository 的靜態概覽:一個 Python 程式碼庫(分析了 45 個檔案),以 whisper 套件為核心,包含 model、decoding、timing、tokenizer、audio 和 transcribe 模組。偵測到的唯一進入點是 whisper/transcribe.py 裡呼叫 cli() 的 __main__ 守衛。requirements.txt 列出七個執行期依賴,另有七個測試檔案;1388 條靜態關係中有 1233 條仍未解析。

它是什麼

這是一個 Python repository:平台 metadata 記錄的主要語言是 Python,有 160526 個 Python 語言位元組。

GitHub 的描述寫著「Robust Speech Recognition via Large-Scale Weak Supervision」(透過大規模弱監督做到穩健的語音辨識):這是作者自己的用語,不是分析出的事實。

它採用 MIT 授權;記錄到的最新 release 是 v20250625(2025-06-26 發布),metadata 快照算到 109056 個 star。

程式碼放在頂層的 whisper 套件目錄,旁邊是 tests 目錄;分析器的摘要回報分析了 45 個檔案,其中 18 個在 whisper 底下。

它怎麼啟動

偵測到的唯一腳本進入點在 whisper/transcribe.py:

if __name__ == "__main__":
    cli()

靜態的核心流程很短:cli()(第 517-619 行)透過一連串 parser.add_argument 呼叫建立 argparse 解析器,並檢查 torch.cuda.is_available。

從 __main__ 出發的六條有界執行路徑,全都停在這些外部或未解析的邊界,所以除了設定參數之外,什麼也沒有追蹤到。

角色分析把 whisper/transcribe.py 標為 entry/orchestration 候選(3 個呼入、18 個靜態呼出)。

結構

  • 關係總數:1388 條靜態邊:1207 個呼叫、181 個 import;155 條在本地解析,1233 條是外部或未解析。
  • whisper/model.py:22 個呼入/20 個呼出,是連結最多的模組,被標為 service/core 候選。
  • whisper/timing.py(16/9)、whisper/utils.py(14/2)、whisper/tokenizer.py(12/2)、whisper/decoding.py(11/13)、whisper/audio.py(8/5):也都是 service/core 候選。
  • whisper/__init__.py(3/4)和 whisper/normalizers/english.py(3/3)是較小的 service/core 候選;tests/test_timing.py(0 入/6 出)是 orchestration 候選。
  • whisper/transcribe.py(3 入/18 出)是 entry/orchestration 候選。
  • 這個樣貌暗示,常被呼叫的模組(model、timing、utils、tokenizer)構成由 transcribe.py 驅動的函式庫核心:這是從呼叫次數推論的,不是執行時的行為。

依賴與測試

  • 依賴證據:requirements.txt 列出七個執行期依賴:numba、numpy、torch、tqdm、more-itertools、tiktoken、triton(>=2.0.0,唯一的版本下限)。
  • 這個 repository 有兩個清單檔(pyproject.toml、requirements.txt),但只有 requirements 形式的檔案會被解析,所以這次分析看不到 pyproject.toml 的依賴表。
  • 測試:7 個測試或類測試檔案,包括 tests/conftest.py;光是 tests/test_timing.py 就定義了 test_dtw、test_dtw_cuda_equivalence、test_median_filter 和 test_median_filter_equivalence。
  • 這裡的安裝與執行說明並不完整,安裝指令也沒有驗證;安裝前請先看過清單檔。

先讀什麼

  1. README.md:整體定位。
  2. pyproject.toml 和 requirements.txt:預期的安裝方式與依賴。
  3. LICENSE:條款。
  4. whisper/transcribe.py:先讀 transcribe()(第 38-514 行),再讀 cli() 解析器(517-619)。
  5. whisper/__init__.py:load_model()、available_models()、_download()。
  6. whisper/model.py:Whisper 類別,把模型的建構和 transcribe() 連起來。

前三步依照分析器的專案層級重要檔案清單;whisper/transcribe.py 被標為可能的進入點。

這次分析的限制

  • 只有靜態分析:反射、執行期依賴注入、動態 import、monkey-patching、產生的程式碼、框架在執行期的接線和動態分派都沒有解析。
  • 1388 條靜態關係中有 1233 條是外部或未解析。
  • 安裝/執行/建置的推論並不完整,安裝指令也沒有驗證。
  • README 的擷取是逐行進行的,可能抓到程式碼行而不是說明文字,所以來自 README 的陳述屬於作者自述,不是已確定的事實。
  • 描述分析器自身能力的 manifest 欄位,不算在 repository 的證據裡。
  • 沒有任何執行時流程(CLI 轉錄、模型載入、GPU 使用)是從頭到尾確定的。

值得注意的符號

  • whisper/model.py:Whisper(259-352)、AudioEncoder(181-211)、TextDecoder(214-256)、MultiHeadAttention(81-146)、ResidualAttentionBlock(149-178)。
  • whisper/decoding.py:detect_language、DecodingOptions、DecodingResult、GreedyDecoder、BeamSearchDecoder。
  • whisper/audio.py:load_audio、pad_or_trim、mel_filters、log_mel_spectrogram;whisper/tokenizer.py:Tokenizer、get_encoding、get_tokenizer。
  • whisper/timing.py:median_filter、dtw_cpu、dtw_cuda、find_alignment、add_word_timestamps。
  • whisper/utils.py:ResultWriter、WriteTXT、SubtitlesWriter、WriteVTT,另有 format_timestamp 和 str2bool。

AudioEncoder/TextDecoder 的命名暗示音訊輸入與文字輸出之間有 encoder-decoder 的分工:這是結構上的推論,沒有經過執行確認。

這次分析無法確定的事

這些是撰寫模型對分析記錄自己的說明。lim_3、exec_* 或 claim_… 這類識別碼指的是那次分析裡的記錄;陳述與依據表引用的也是同一批記錄。

  • 1388 條靜態關係中有 1233 條是外部或未解析,所以大部分跨模組的行為沒有被追蹤。
  • pyproject.toml 的依賴表沒有被解析,所以依賴的全貌可能不完整。
  • 沒有執行任何程式碼;CLI 的行為、模型載入和 CUDA 路徑都沒有驗證。
  • 模組角色來自靜態呼叫次數和是否屬於進入點,所以它們是結構上的線索,不是量測出的執行時角色。
  • 被標記為重要檔案的 readme.md 被排除了,因為這個路徑不在被分析的檔案清單裡(大小寫不同的比對)。
陳述與依據 — 36 條陳述,36 條經獨立驗證者確認(英文原文)

上面每一句都是一條陳述,各自引用分析裡的依據編號。編號是這次分析內部的;表格列出每條陳述靠的是什麼、依據有多強。

Every substantive statement above is a claim bound to grounding IDs of the analyzed revision. Global grounding IDs are namespaced by the analysis run.

Claim Epistemic status Verifier Grounding (global IDs)
c1 Platform metadata records Python as the primary language, with {"Python": 160526} language bytes. observed supported analysis_ee52b965e66bce7b:meta_primary_language, analysis_ee52b965e66bce7b:meta_language_bytes
c2 The repository's GitHub description is "Robust Speech Recognition via Large-Scale Weak Supervision". author_claimed supported analysis_ee52b965e66bce7b:meta_description
c3 The project is MIT licensed, and its latest recorded release is v20250625, published 2025-06-26. observed supported analysis_ee52b965e66bce7b:meta_license, analysis_ee52b965e66bce7b:meta_latest_release, analysis_ee52b965e66bce7b:important_5
c4 Source code sits in a top-level whisper package directory alongside a tests directory. observed supported analysis_ee52b965e66bce7b:subsys_2, analysis_ee52b965e66bce7b:subsys_1
c5 The analyzer's summary reports 45 analyzed files, 18 of them in the whisper directory and 7 in tests. inferred supported analysis_ee52b965e66bce7b:summary_repository
c6 The repository had 109056 stargazers as of 2026-09-14 per platform metadata. observed supported analysis_ee52b965e66bce7b:meta_stars
c7 whisper/transcribe.py contains a __main__ execution guard whose body calls cli(); cli() spans lines 517-619 of that file. observed supported analysis_ee52b965e66bce7b:py_entry_e9d8e13fe4d4, analysis_ee52b965e66bce7b:py_func_a31ce0a9065c
c8 Six bounded static execution paths start at whisper/transcribe.py:__main__ and pass through cli(); each terminates at an external or unresolved call — argparse.ArgumentParser, parser.add_argument (repeatedly), or torch.cuda.is_available. inferred supported analysis_ee52b965e66bce7b:exec_210328c6174a, analysis_ee52b965e66bce7b:exec_873b5bbd14cf, analysis_ee52b965e66bce7b:exec_deca1b3baebb, analysis_ee52b965e66bce7b:exec_210dbd98fcf8, analysis_ee52b965e66bce7b:exec_c8ba6b2b53c8, analysis_ee52b965e66bce7b:exec_bd66f158fc76
c9 The analyzer labels whisper/transcribe.py an entry/orchestration candidate, with 3 incoming and 18 outgoing static calls. inferred supported analysis_ee52b965e66bce7b:role_16
c10 The analyzer recorded 1388 static relations (1207 calls, 181 imports); 155 resolved locally and 1233 stayed external or unresolved. observed supported analysis_ee52b965e66bce7b:relation_counts
c11 whisper/model.py is the most connected module (22 incoming, 20 outgoing calls) and is labeled a service/core candidate. inferred supported analysis_ee52b965e66bce7b:role_10
c12 whisper/decoding.py (11/13), whisper/timing.py (16/9), whisper/tokenizer.py (12/2), whisper/utils.py (14/2), and whisper/audio.py (8/5) are each labeled service/core candidates. inferred supported analysis_ee52b965e66bce7b:role_9, analysis_ee52b965e66bce7b:role_14, analysis_ee52b965e66bce7b:role_15, analysis_ee52b965e66bce7b:role_18, analysis_ee52b965e66bce7b:role_8
c13 whisper/normalizers/english.py (3/3) and whisper/__init__.py (3/4) are smaller service/core candidates; tests/test_timing.py (0 incoming, 6 outgoing) is an orchestration candidate. inferred supported analysis_ee52b965e66bce7b:role_13, analysis_ee52b965e66bce7b:role_6, analysis_ee52b965e66bce7b:role_3
c14 whisper/transcribe.py is labeled an entry/orchestration candidate (3 in / 18 out); combined with the degree numbers, this suggests it drives a core of often-called modules — an inference from call degree, not runtime behavior. inferred supported analysis_ee52b965e66bce7b:role_16, analysis_ee52b965e66bce7b:role_10, analysis_ee52b965e66bce7b:role_14, analysis_ee52b965e66bce7b:role_18, analysis_ee52b965e66bce7b:role_15
c15 requirements.txt lists seven runtime dependencies: numba, numpy, torch, tqdm, more-itertools, tiktoken, and triton (>=2.0.0, the only one with a version floor). observed supported analysis_ee52b965e66bce7b:dep_1, analysis_ee52b965e66bce7b:dep_2, analysis_ee52b965e66bce7b:dep_3, analysis_ee52b965e66bce7b:dep_4, analysis_ee52b965e66bce7b:dep_5, analysis_ee52b965e66bce7b:dep_6, analysis_ee52b965e66bce7b:dep_7
c16 The repository contains 2 build/dependency manifest files, but only requirements-style manifests are parsed; pyproject.toml dependency tables are not. observed supported analysis_ee52b965e66bce7b:ev_manifest_1, analysis_ee52b965e66bce7b:lim_3
c17 The tests directory contains 7 test files or test-like files, including tests/conftest.py. observed supported analysis_ee52b965e66bce7b:ev_tests_1
c18 tests/test_timing.py defines four selected test functions: test_dtw, test_dtw_cuda_equivalence, test_median_filter, and test_median_filter_equivalence. observed supported analysis_ee52b965e66bce7b:py_func_6b8d72c7960f, analysis_ee52b965e66bce7b:py_func_681d1ea53357, analysis_ee52b965e66bce7b:py_func_cf597b3d5d25, analysis_ee52b965e66bce7b:py_func_59835aca4f63
c19 Install guidance from the analyzer is partial: installation commands are not verified. observed supported analysis_ee52b965e66bce7b:lim_2
c20 The packet marks five important files: README.md, pyproject.toml, requirements.txt, LICENSE, and whisper/transcribe.py (flagged as a likely entrypoint). observed supported analysis_ee52b965e66bce7b:important_1, analysis_ee52b965e66bce7b:important_3, analysis_ee52b965e66bce7b:important_4, analysis_ee52b965e66bce7b:important_5, analysis_ee52b965e66bce7b:important_6
c21 A grounded reading order: README.md first for orientation, then pyproject.toml and requirements.txt for setup, then LICENSE, then whisper/transcribe.py. inferred supported analysis_ee52b965e66bce7b:important_1, analysis_ee52b965e66bce7b:important_3, analysis_ee52b965e66bce7b:important_4, analysis_ee52b965e66bce7b:important_5, analysis_ee52b965e66bce7b:important_6
c22 whisper/transcribe.py defines transcribe() at lines 38-514 and cli() at lines 517-619. observed supported analysis_ee52b965e66bce7b:py_func_ed32c16d5c96, analysis_ee52b965e66bce7b:py_func_a31ce0a9065c
c23 whisper/__init__.py defines load_model(), available_models(), and _download(). observed supported analysis_ee52b965e66bce7b:py_func_3caacfc5e3d9, analysis_ee52b965e66bce7b:py_func_7ca1feb87e67, analysis_ee52b965e66bce7b:py_func_a5de7f0ce3d2
c24 After transcribe.py, read the Whisper class in whisper/model.py and load_model() in whisper/__init__.py to connect model construction with the transcription flow. inferred supported analysis_ee52b965e66bce7b:py_class_258632aeaf38, analysis_ee52b965e66bce7b:py_func_3caacfc5e3d9
c25 The analyzer documents that reflection, runtime dependency injection, dynamic imports, monkey-patching, generated code, framework runtime wiring, and dynamic dispatch are not resolved. observed supported analysis_ee52b965e66bce7b:lim_1
c26 1233 of 1388 static relations remain external or unresolved. observed supported analysis_ee52b965e66bce7b:relation_counts
c27 Install/run/build inference is partial and installation commands are not verified. observed supported analysis_ee52b965e66bce7b:lim_2
c28 README claim extraction is line-based and may capture code lines instead of prose, so README-derived statements are author-claimed, not established. observed supported analysis_ee52b965e66bce7b:lim_4
c29 Manifest fields such as project.actual_capabilities describe the analyzer itself, not the analyzed repository, and are excluded from repository evidence. observed supported analysis_ee52b965e66bce7b:lim_5
c30 End-to-end runtime behavior is not established: the packet itself marks dynamic dispatch, reflection, generated-code, and runtime framework behavior as unresolvable statically. unresolved supported analysis_ee52b965e66bce7b:claim_7c9925f9c8f4
c31 whisper/model.py defines Whisper (lines 259-352), AudioEncoder (181-211), TextDecoder (214-256), MultiHeadAttention (81-146), and ResidualAttentionBlock (149-178). observed supported analysis_ee52b965e66bce7b:py_class_258632aeaf38, analysis_ee52b965e66bce7b:py_class_b562f875502f, analysis_ee52b965e66bce7b:py_class_785260aa3f04, analysis_ee52b965e66bce7b:py_class_4ed9b8eada83, analysis_ee52b965e66bce7b:py_class_52b8e7fc7bc9
c32 whisper/decoding.py defines detect_language(), DecodingOptions, DecodingResult, GreedyDecoder, and BeamSearchDecoder, among other classes. observed supported analysis_ee52b965e66bce7b:py_func_aa60011d71a2, analysis_ee52b965e66bce7b:py_class_73648de04931, analysis_ee52b965e66bce7b:py_class_aa1a4e6d4813, analysis_ee52b965e66bce7b:py_class_f015fc72fa5b, analysis_ee52b965e66bce7b:py_class_61d8ff25d720
c33 whisper/audio.py defines load_audio(), pad_or_trim(), mel_filters(), and log_mel_spectrogram(); whisper/tokenizer.py defines Tokenizer, get_encoding(), and get_tokenizer(). observed supported analysis_ee52b965e66bce7b:py_func_0ac35ae1bbef, analysis_ee52b965e66bce7b:py_func_915c49ac6d81, analysis_ee52b965e66bce7b:py_func_499661f28292, analysis_ee52b965e66bce7b:py_func_d200eb51a4b3, analysis_ee52b965e66bce7b:py_class_814237ef1902, analysis_ee52b965e66bce7b:py_func_0d16d45ecbed, analysis_ee52b965e66bce7b:py_func_5be6d8888ea5
c34 whisper/timing.py defines median_filter(), dtw_cpu(), dtw_cuda(), find_alignment(), and add_word_timestamps(). observed supported analysis_ee52b965e66bce7b:py_func_dbcb40294500, analysis_ee52b965e66bce7b:py_func_56b7a739b461, analysis_ee52b965e66bce7b:py_func_fe8c28e8968b, analysis_ee52b965e66bce7b:py_func_e1372f4b817b, analysis_ee52b965e66bce7b:py_func_69d75b442d80
c35 whisper/utils.py defines the writers ResultWriter, WriteTXT, SubtitlesWriter, and WriteVTT, plus helpers such as format_timestamp() and str2bool(). observed supported analysis_ee52b965e66bce7b:py_class_892b197925de, analysis_ee52b965e66bce7b:py_class_16f1ca58bef3, analysis_ee52b965e66bce7b:py_class_e87997a189d5, analysis_ee52b965e66bce7b:py_class_610f01bd8eb9, analysis_ee52b965e66bce7b:py_func_8c9c9c69899c, analysis_ee52b965e66bce7b:py_func_24929df50c11
c36 The AudioEncoder/TextDecoder pair in whisper/model.py suggests an encoder-decoder split between audio input and text output; this is a structural reconstruction, not confirmed by execution. inferred supported analysis_ee52b965e66bce7b:architecture_reconstruction, analysis_ee52b965e66bce7b:py_class_b562f875502f, analysis_ee52b965e66bce7b:py_class_785260aa3f04

來源、權利與說明 · attribution-license-templates/v0.1

權利聲明。原始 repository 託管於 GitHub。其原始碼、文件、名稱、媒體與相關素材的權利,仍屬各自的作者、貢獻者與其他權利人所有,並受該 repository 的授權條款約束。

Original repository hosted on GitHub. Repository source code, documentation, names, media, and related project materials remain subject to the rights of their respective authors, contributors, and other rights holders and to applicable repository license terms.

平台聲明。GitHub 是所連結 repository 的來源託管平台。GitHub 與相關標誌為 GitHub, Inc. 的商標。除非另有明確說明,EVEMISS Technology 與 GitHub 之間沒有隸屬或背書關係。

GitHub is the source hosting platform for the linked repository. GitHub and related marks are trademarks of GitHub, Inc. EVEMISS Technology is not affiliated with or endorsed by GitHub unless explicitly stated otherwise.

這一頁怎麼來的。本頁由對應版本的 repository 分析與 AI 輔助的編輯工具產生。技術陳述綁定所分析的版本;原始 repository 更新後,可能重新驗證。

This page was produced using revision-aware repository analysis and AI-assisted editorial tooling. Technical claims are tied to the analyzed repository revision and may be revalidated when the source repository changes.

本頁包含 AI 輔助分析,並經 EVEMISS Technology 人工與 AI 協力審閱。 · 回報權利疑慮 · 所有 repository 指南