Training-data sheet · snapshot 2026-09-24 07:20 EDT (v4 complete, v5 data building)Sichos / asr · prepared for the training scientist
Export: this page prints cleanly (File → Print → Save as PDF; the layout has print styles). A PDF, the standalone HTML, every chart as SVG and PNG, and the underlying tables as CSV are in asr/reports/corpus-datasheet-export.zip.
Rebbe Yiddish ASR Corpus
Everything the speech-recognition models in this project train and test on, updated 2026-09-24 with the new source data, the v4 results, the audit of the delivered transcripts and the plan for v5: where the audio and text come from, how they were paired and cut into clips, what was filtered, how the fixed evaluation sets were chosen, the pseudo-labeled extension, and what each training run did with it. All figures are computed from the manifests in asr/data/manifests/ and the evaluation reports in asr/reports/; nothing here is estimated.
1,468paired days (audio + Yiddish hanacha)
30,212clips ≤ 28 s, 168 h
1.25 Mreference words
27,833clean training clips, 153 h
20 / 20fixed dev / test days
+1,758paired days arriving in v5 (5775–5781): 306 h, 2.31 M words
Hours are audio inside clips. Words are counted on the training-target text (punctuation kept, vowel points removed). Hebrew years: 5711 = 1950/51 … 5752 = 1991/92; 5783 = 2022/23 … 5787 = 2026/27.
What one training example is
A clip is a stretch of at most 28 s of the Rebbe's speech, cut from a Daily Sicha recording, paired with the words the published Yiddish hanacha (the written transcript of the talk) gives for that stretch. The model sees the 16 kHz mono audio and learns to emit the text. A typical clip is 20 s long and carries about 42 words at 2.1 words per second.
id 5783/001/000 · 38.22–54.59 s · 16.4 s · timing: site sync (offset-corrected)
[the example text is withheld on the public page: the hanachos belong to The Daily Sicha and are not republished here]
One real training target. Note the conventions the model must learn: Yiddish in unpointed Hebrew script, loshon-kodesh phrases embedded verbatim, source names abbreviated with gershayim (רמב"ן, של"ה), quoted terms in quotation marks.
Where the data comes from
The Daily Sicha (thedailysicha.com) publishes one ~10-minute excerpt of a farbrengen per day with a Yiddish hanacha, and for recent years a sync file with segment timings. The project holds the daily mp3s for five publication years (5783–5787) and the site's text for each; the source talks span 1950 to 1992. Four older archive years (5778–5781, 1,005 files) have audio only and are the transcription target of the project.
Pair daysResolve each mp3 to its Hebrew date, fetch the day's hanacha and sync JSON.fetch_days.py
Measure offsets292 of 613 daily files carry a dedication before the excerpt (median 9.2 s); cross-correlation finds where the excerpt starts.compute_offsets.py
Align text to audioWhere the site has no timings, a Yiddish Whisper model transcribes with word times and the hypothesis is anchored to the reference (exact, then fuzzy).align_day.py
Cut clipsMerge consecutive timed segments up to 28 s; write FLAC clips and manifests; split by day.build_dataset.py
FilterDrop clips whose baseline transcription disagrees wildly with the reference (misaligned timings).filter_manifest.py
ExtendTranscribe the 1,005 text-less files with the best model and turn confident windows into pseudo-labeled clips.build_pseudo_dataset.py
Clips by Daily Sicha publication year. 1,460 of the 1,468 paired days yielded clips (8 days failed alignment). Per year: 5783 = 242 days, 5784 = 318, 5785 = 290, 5786 = 291, 5787 = 319.How each clip got its timings. 27,912 clips (92%) are timed by the project's own anchor alignment, 2,300 (8%) by the site's sync files corrected for the dedication offset. The test set is deliberately all site-timed; dev is all self-aligned (see splits).
Hours of paired audio by year of the original farbrengen (5711–5752; 637 clips from 32 days have no recorded year and are omitted here). The bulk of the material is 5734–5749; the shaded region marks 5739 onward, the years the requested transcriptions must cover. Older recordings are noisier and the speaker is younger, so the model sees both conditions.
New since 2026-09-23: hanachos and audio for 5775–5781
The Daily Sicha delivered the Yiddish hanachos as one PDF per year (5768–5782 and 5784–5787) together with the recordings of 5775–5777. The four years the project had been transcribing without text (5778–5781) and three more years are now paired: 1,758 days, 306 h of recordings, 2.31 M words, more than doubling the corpus once aligned. Nine PDFs have a clean text layer. Ten (5768–5770, 5774–5780) are set in custom-encoded fonts: every glyph code is a fixed permutation of the Hebrew alphabet, different per font. They are decoded exactly rather than OCR'd: pipeline/pdf_fontmap.py hill-climbs each font's code→letter map by the share of decoded words that exist in the training vocabulary (0.91–0.93 for the body font) and pins the bold header font with the words every day's header shares (בס"ד. "השיחה היומית" ליום, the year, התוכן, the month names). Tesseract on the same pages scored 0.39 WER; the exact decode leaves no systematic errors (the drafts-vs-text numbers below equal the clean-layer year).
Decode the fontsLearn one code→letter map per font from ~100k sample words; bold font pinned by the header crib.pdf_fontmap.py
Split into daysHeader line → day; the printed [NNN] number → the mp3 stem. All 1,758 days matched, none missing or duplicated.pdf_hanachos.py
Day recordsSame record shape as the site years, text as one paragraph per PDF line, no site timings (offset 0).build_pdf_days.py
Align on the fleet7 Mac Studios × 5 workers, production CT2 model anchors the text to the audio (anchor rates 0.94–0.98).align_day.py
Build + scanClips ≤ 28 s; the production model transcribes every train clip; clips with CER > 0.6 are dropped (per-year drop rate = text-quality signal).build_dataset.py · filter_manifest.py
new years (aligning now)v4 training years
Paired days per Daily Sicha publication year. Hours are recording time before clipping (roughly 65% of it ends up inside clips). Text for 5768–5770, 5774 and 5782 exists too (about 1.4 M words) but no audio is in hand, so those years are text-only for now.
Three traps in the PDFs, all fixed at the glyph level. (1) 5776–5778 carry a diagonal watermark (הנחה פרטית בלתי מוגה, 80–85 pt): its letters landed on body rows and fused words (2,100–2,700 fourteen-letter "words" per year against ~130 in clean years) and left ~500 stray single letters per year; glyphs over 40 pt are now dropped before rows are formed. (2) Every page's footer small print (contact line, 6–10 pt, Arial/Tahoma) decoded through the wrong font into junk tokens, up to 2,788 per year; rows are now kept only in the body font family, not by size (two 10 pt days would have been emptied by a size cut), and the training-text normalizer drops any token with Latin letters. (3) The bold header font learned without the crib landed in a wrong permutation on 5775/5776 (0.03 in-vocabulary, zero days found); with the crib it reaches 0.82–0.85. After the fixes the 5778 drafts score 0.072 against the text instead of 0.120: the gap was the text, not the audio.
Splits
Splitting is by day, never by clip, so no test audio shares a recording with training audio. The rule, fixed on 2026-09-16 and never changed:
test: 20 evenly spaced days among the site-timed days (Tamuz 5786-Tishrei 5787) whose source farbrengen is 5739 or later and whose offset is known; dev: 20 evenly spaced days among the remaining text-only days with known offsets. Fixed on 2026-09-16; never train on these days.
Caveats for the scientist: (1) the dev WER printed during training uses a fixed 200-clip subset of dev decoded with the HF generate path; the full-dev and test numbers in the results section use the full sets. (2) One test day, 9 Tishrei 5787, scores 0.28–0.34 for every model and is a suspected reference-quality outlier; it is kept because the split is frozen. (3) Out-of-vocabulary rate against the clean training text: dev 1.1%, test 1.2%.
Clip anatomy
Clip duration. Mean 20.0 s, median 21.2 s, 10th–90th percentile 11.1–26.8 s, max 28.3 s. The builder merges consecutive timed segments until the next one would exceed 28 s, so most clips sit just under the Whisper 30 s window.Words per clip. Mean 41.5, median 43; 2.07 words per second of audio overall. Clips per day: median 21 (1–57).
Alignment quality
For the 92% of clips timed by the project's own aligner, the anchor rate is the share of a clip's reference words that were matched exactly or fuzzily to a word the Yiddish Whisper model heard at a known time; the remaining words are placed by interpolation. Clips below 0.30 are excluded at build time, and clip boundaries must fall on anchored words.
Anchor rate of self-aligned clips (n = 27,912; mean 0.65, median 0.66). Validation against the site's own timings on 3 Tishrei 5787 (113 segments): median start difference 0.33 s, 90th percentile 0.95 s, 83% of segments within 1 s at both ends.
Cleaning
The whole train split was transcribed once with the untouched Yiddish model and each clip's hypothesis compared with its reference. Clips with character error rate above 0.60, or a hypothesis/reference length ratio outside 0.5–2.0, were removed: 1,440 clips (4.9%), 1,436 of them by the CER rule (median CER of the dropped clips 0.75). These are almost always timing errors, where the text does not belong to the audio, not hard speech. The v0 model trained on the unfiltered set; every run since uses train.clean.jsonl.
Vocabulary
Clean training text: 1,161,335 tokens, 27,712 distinct words (letters only, no marks), 10,646 of them seen once. The 1,000 most frequent words cover 81.8% of tokens, the top 5,000 cover 94.3%. Words seen 100+ times are 4.0% of the types but 82.7% of the tokens — the long tail of names and sources is where the remaining errors live.
Text conventions and normalization
Training target
normalize.training_text: the hanacha paragraph text with vowel points (nikkud) removed, punctuation and the abbreviation marks kept (geresh ' and gershayim "), bracket characters removed but the bracketed words kept (they are spoken asides; dropping them raised WER for every model). 27,644 of 30,212 raw references carried nikkud; 68.8% of clips contain at least one gershayim abbreviation.
Scoring text
normalize.scoring_text: Hebrew letters only. All WER and CER figures anywhere in the project are computed on this form, so punctuation, quotes and spelling of marks never count as errors, but a misspelled abbreviation does.
Reference frames
Site sync timings can refer to the raw excerpt or to the daily file with its dedication; calibrate_site_timings.py decides per day and compute_offsets.py supplies the offset. Days whose offset could not be established are never used with site timings.
Decoder note
Whisper's default token-suppression list blocks the ASCII double quote, which in Hebrew text is the gershayim. Every fine-tuned model learned to write רש"י and הקב"ה but could not emit them through faster-whisper until 2026-09-21; the production decoder (HF fp32 on the GPU, VAD windows, no timestamp tokens) has no such list.
Pseudo-labeled extension (v3 self-training)
Superseded for training on 2026-09-23: the 5778–5781 recordings now have their real hanachos (section above), so v5 trains on human text for these years; the pseudo-labels remain as a documented experiment. Result of self-training (v3 + v4): it lifts weaker starting points by 0.5–1 point but never beats the best supervised recipe on the same data (production 0.081 → 0.081; large lr 2e-5 0.082 → 0.085).
A second pseudo-label set (26,797 clips, 169.4 h) was cut on 2026-09-22 from the third transcription pass (model turbo lr 3e-5, whole-file WER 0.081); the two v4 self-training runs use it, the v3 runs below used the first set. Result of v3 on the full test set: 0.082 / 0.084 / 0.086 whole-file for the three recipes — each 0.5–1 point better than its starting model, none better than the supervised lr 3e-5 model (0.081).
The 1,005 recordings of 5778–5781 have no human text. They were transcribed with the best model (turbo full fine-tune, lr 2e-5) through the production decoder; each VAD window of up to 28 s became a clip with the model's text as its label. Windows with mean token log-probability below −0.6, fewer than two words, or repetitive text (decoder loops, 32 windows) were dropped, and labels were normalized exactly like human targets.
Pseudo-labeled clips by archive folder. 26,778 clips, 169.3 h, mean 22.8 s, 77,201 gershayim in the labels. Confidence (mean token log-prob per window): 10th percentile −0.065, median −0.035, 90th −0.020.
Mixed training set for v3
train.clean_pseudo.jsonl = 27,833 human-timed clips + 26,778 pseudo-labeled clips = 54,611 rows, about 322 h. At batch 8 × 4 accumulation that is 1,707 steps per epoch. The pseudo-labels inherit the labeling model's error rate (about 9% of words), so v3 tests whether in-domain audio with imperfect labels beats the same continuation on real data alone (run on bigmac03 in v2 for exactly this comparison).
Rows carry timing_source: pseudo, min_logprob and n_segments, so they can be filtered or down-weighted later without rebuilding.
Training runs
All runs fine-tune ivrit.ai's Yiddish-adapted Whisper (large-v3 or large-v3-turbo) in fp32 on Apple-silicon GPUs, one run per Mac Studio, from the same clean manifest; batch 8 × 4 accumulation (32 clips per step), 870 steps per epoch, linear decay with 200 warm-up steps, evaluation every 500 steps on the 200-clip dev subset. Lines are dev WER; a dot is an evaluation.
turbo · full FT · lr 2e-5 · 3 ep 2610 steps · last dev 0.0765
turbo · full FT · lr 1e-5 · 3 ep 2610 steps · last dev 0.0793
turbo · LoRA r64 · 3 ep 2610 steps · last dev 0.0836
turbo · frozen encoder · lr 2e-5 · 4 ep 3480 steps · last dev 0.1959
turbo · frozen encoder · lr 1e-5 · 4 ep 3480 steps · last dev 0.2277
large-v3 · full FT · lr 1e-5 · 1 ep 870 steps · last dev 0.1025
large-v3 · LoRA r64 · 1 ep 870 steps · last dev 0.1029
large-v3 · frozen encoder · 2 ep 1740 steps · last dev 0.1664
continue best v1 (turbo lr 2e-5) · lr 1e-5 · 3 ep 2610 steps · last dev 0.0805
turbo · full FT · lr 3e-5 · 4 ep 3480 steps · last dev 0.0730
large-v3 · full FT · lr 1e-5 · 3 ep 2610 steps · last dev 0.0809
large-v3 · full FT · lr 2e-5 · 3 ep 2610 steps · last dev 0.0723
continue v1 lr 2e-5 on real + pseudo · lr 1e-5 · 2 ep 3414 steps · last dev 0.0708
continue v1 lr 1e-5 on real + pseudo · lr 1e-5 · 2 ep 3414 steps · last dev 0.0693
fresh turbo on real + pseudo · lr 2e-5 · 2 ep 3414 steps · last dev 0.0686
v4: production model self-trained on pass-3 labels · lr 1e-5 · 2 ep 3416 steps · last dev 0.0699
v4: large lr 2e-5 + pass-3 pseudo · lr 1e-5 · 2 ep 3416 steps · last dev 0.0733
turbo · full FT · lr 4e-5 · 4 ep 3480 steps · last dev 0.0847
turbo · full FT · lr 5e-5 · 4 ep 3480 steps · last dev 0.0713
large-v3 · full FT · lr 3e-5 · 4 ep 3480 steps · last dev 0.0751
turbo · full FT · lr 3e-5 · 4 ep · seed 7 (repeat) 3480 steps · last dev 0.0740
Reading the curves: freezing the encoder (amber) plateaus at 0.19–0.22, LoRA (teal, purple) lands near 0.084, full fine-tuning of the turbo model reaches 0.0765 at lr 2e-5. v2 finished with large-v3 at lr 2e-5 on 0.0723 and turbo at lr 3e-5 on 0.0730 (both better than the v1 winner); continuing the v1 winner at lr 1e-5 only reached 0.0757. The v3 self-training runs ended with the best in-training scores of all (0.0686–0.0693) but scored 0.082–0.086 whole-file on the full test set, behind the production model — the 200-clip dev subset overstates models trained on pseudo-labels, so only the full test set ranks them. Pseudo-labels lifted each starting model by 0.5–1 point without surpassing the best supervised recipe. v4 probes higher learning rates, the turbo recipe on large-v3, a seed repeat of the production run, and self-training from the strongest starting points with fresher labels.
Results on the fixed test set
Two metrics, both against the human hanacha of the 20 test days: clip WER on the 488 test clips (audio cut on known timings) and whole-file WER where each day's full recording is decoded with voice-activity chunking and scored against the whole hanacha — the number that matters for the product. Decoder matters as much as model: the same weights score 0.132, 0.124 and 0.088 whole-file under the three decoders tried. The v2 winner on whole files is the turbo model trained at lr 3e-5 for 4 epochs (0.081, CER 0.045); large-v3 at lr 2e-5 is best on clips (0.083) but decodes three times slower and is a hair worse on whole files (0.082).
clip test WERwhole-file test WER
Baselines are the ivrit.ai models without any fine-tuning on this corpus. "CT2 int8" = CTranslate2/faster-whisper on CPU; "quotes freed" = same with the gershayim token un-suppressed; "GPU path" = HF fp32 model on MPS with VAD windows and no timestamp tokens (pipeline/hf_longform.py). Acceptance targets were ≤ 0.30 clip and ≤ 0.35 whole-file.
model · decoder
clip dev WER
clip test WER
clip test CER
whole-file WER
whole-file CER
v2: turbo full lr 3e-5, 4 ep · GPU path (production since 09-22)
0.080
0.087
0.050
0.081
0.045
production model · auto windows (decoder default since 09-23)
–
–
–
0.081
0.045
v4: large-v3 full lr 3e-5, 4 ep (resumed after the disk crash) · GPU path
0.080
0.086
0.051
0.081
0.047
v4: production recipe, seed 7 (repeat) · GPU path
0.081
0.087
0.050
0.080
0.045
v4: production model self-trained on pass-3 labels · GPU path
0.079
0.087
0.050
0.081
0.045
v4: turbo full lr 5e-5, 4 ep · GPU path
0.079
0.089
0.053
0.082
0.046
v4: turbo full lr 4e-5, 4 ep · GPU path
0.084
0.090
0.050
0.088
0.048
v4: large-v3 lr 2e-5 model + pass-3 pseudo-labels · GPU path
0.080
0.089
0.050
0.085
0.047
v2: large-v3 full lr 2e-5, 3 ep · GPU path
0.077
0.083
0.048
0.082
0.047
v3: turbo lr 1e-5 model + 2 ep on real + pseudo · GPU path
0.079
0.090
0.051
0.082
0.046
v3: turbo lr 2e-5 model + 2 ep on real + pseudo · GPU path
0.080
0.090
0.052
0.084
0.046
v3: fresh turbo lr 2e-5, 2 ep on real + pseudo · GPU path
0.079
0.092
0.053
0.086
0.047
v2: continue turbo lr 2e-5 at lr 1e-5, 3 ep · GPU path
0.082
0.091
0.051
0.086
0.048
v2: large-v3 full lr 1e-5, 3 ep · GPU path
0.083
0.091
0.051
0.087
0.048
turbo full lr 2e-5 · GPU path (production 09-21)
0.080
0.093
0.052
0.088
0.048
turbo full lr 1e-5 · GPU path
0.088
0.098
0.053
0.093
0.050
turbo full lr 2e-5 · CT2 int8, quotes freed
0.082
0.095
0.053
0.124
0.081
turbo full lr 1e-5 · CT2 int8, quotes freed
0.093
0.101
0.057
0.121
0.075
turbo full lr 2e-5 · CT2 int8, default suppression
0.092
0.105
0.057
0.132
0.084
turbo full lr 1e-5 · CT2 default
0.100
0.110
0.060
0.128
0.076
turbo LoRA r64 · CT2 default
0.101
0.118
0.064
0.127
0.073
large-v3 full lr 1e-5, 1 ep · CT2 default
0.110
0.126
0.068
0.144
0.083
large-v3 LoRA r64, 1 ep · CT2 default
0.115
0.134
0.071
0.147
0.084
large-v3 frozen encoder · CT2 default
0.157
0.182
0.099
0.199
0.117
turbo frozen encoder lr 2e-5 · CT2 default
0.195
0.225
0.119
0.239
0.138
turbo frozen encoder lr 1e-5 · CT2 default
0.212
0.244
0.128
0.266
0.148
v0: frozen encoder, 1,000 steps, unfiltered data
0.254
0.284
0.147
0.305
0.168
baseline ivrit-ai yi-whisper-large-v3 (no fine-tuning)
0.496
0.531
0.274
0.534
0.275
baseline ivrit-ai yi-whisper-large-v3-turbo (no fine-tuning)
0.542
0.571
0.286
0.569
0.283
Accuracy on the product years: the delivered transcripts against the hanachos
With the PDFs in hand, the 1,005 machine transcripts of 5778–5781 (third pass, production model) can be scored against the printed text for the first time. These are whole-file numbers over all days, including days whose printed text is a different sicha or a heavily edited one, so they are upper bounds; the fixed test set (0.081) sits in the same range.
year
days
reference words
WER
CER
median day WER
reference text
5778
247
319,404
0.072
0.041
0.059
decoded PDF (David + bold crib)
5779
266
341,973
0.070
0.040
0.057
decoded PDF
5780
245
306,533
0.073
0.042
0.057
decoded PDF
5781
247
312,740
0.090
0.050
0.071
clean text layer; 0.080 without the 2 mismatched days
A decoder failure the earlier audit missed. Comparing word counts exposed 14 transcripts far shorter than their text: the silero voice-activity detector had kept 3 windows of a 638 s loud, clipped recording (31 words for 011 18-Tishrei 5781) and dropped 11–52% of the audio in 13 more files, while the "0 empty transcripts" check passed. Lower VAD thresholds recover little; decoding every gap as loud as the speech recovers it but hallucinates on gaps that are crowd noise or singing. The decoder now (1) falls back to fixed windows when the VAD covers under 30% of a recording and (2) by default adds loud gaps as extra windows and keeps each only when its decode is confident (mean token log-prob above −0.35 and at least 1 word/s). Validation: identical output on normal files (control set 0.0955 = 0.0955; the 20-day test set 0.0814 vs 0.0814), 0.231 → 0.135 on the 13 affected files. The ten truncated transcripts were replaced; the originals are kept in output_superseded/.
Where the errors are now
v1 turbo lr 1e-5, CT2 default suppression (test 0.110)v2 turbo lr 3e-5, GPU path (test 0.087)
Error rate of reference words by how often the word occurs in the clean training text (substitutions + deletions on 42,368 dev+test words), plus abbreviations written with gershayim. Frequent words are right 95% of the time; words never seen in training are wrong half the time. The abbreviation class fell from 30% to 17.5% once the decoder could emit the quote and the model improved; what remains there is mostly abbreviation-vs-spelled-out variation that the references themselves are inconsistent about (עאכו"כ vs על אחת כמה וכמה, בנוגע vs בהנוגע). The single largest confusion is דער/די/דעם (gender and case of the article).
v5: the enlarged dataset and the runs on it (in progress)
Status at the snapshot: the fleet has aligned 5781 and is at 1,170 of the remaining 1,511 days (5775–5780, wave 2; ~6 days per minute). When it finishes, every node and this Mac rebuild the dataset from the same day records, the production model scans the train split, the filter drops clips with CER > 0.6, and a guarded launcher starts the runs below unless the cleaned manifest is small or a year lost more than 25% of its clips. Rules kept: data/splits.json unchanged (the 20 dev and 20 test days stay in 5786–5787); the new years go to train only, except test_b = 20 evenly spaced 5781 days (clean text layer) that are never trained on and are reported next to the main test set from now on. Expected size: about 66,000 clips, ~370 h (30,212 clips / 168 h in v4).
run
node
recipe
why
expected
turbo-full-v5-cont-lr1e5-2ep
bigmac01
production model continued 2 ep on all data, lr 1e-5
cheapest read of what the new years add
~12 h
turbo-full-v5-lr3e5-2ep
bigmac02
production recipe from the base on all data, 2 ep, lr 3e-5
same optimizer steps as the 4-epoch v4 run on 2.5× the clips
~12 h
large-full-v5-lr2e5-2ep
supermac01 (512 GB)
large-v3 from the base on all data, 2 ep, lr 2e-5
best clip model in v4 (0.083); the target model
~35 h on a Mac Studio; hours on the GB10 with CUDA/bf16
v4 verdict, complete (all six runs scored): on the old data nothing beats the production recipe; the seed repeat (0.080 vs 0.081 whole-file) puts run-to-run noise at ~0.001, so table differences under 0.003 mean nothing. Higher learning rates (4e-5, 5e-5) are worse; self-training never wins; large-v3 at lr 3e-5 (0.086 clips / 0.081 whole-file) equals the turbo on whole files at 2.5× the cost and does not reach large-v3 at lr 2e-5 on clips. The v5 data is the only remaining lever.
Next steps
Finish the v5 data (today). Alignment wave 2 → dataset → production-model scan → filter. Read the per-year drop rate: a year above ~10% means a decode or alignment problem in that year's text, not bad audio.
Train the three v5 runs (Thu–Sat) and score each on the fixed test set through the GPU path (whole-file WER is the ranking metric) and on test_b. Decide by whole-file test WER; anything under 0.003 apart is a tie.
If a v5 model wins: a fourth Phase 7 pass with it and the auto windows over 5775–5781 (the project owner's go-ahead needed; the third pass stands until then), then the per-year drafts-vs-hanacha table again.
Attack the remaining errors where they are: half of the remaining word errors are words never seen in training (names, sources, loshon-kodesh); the text-only years (5768–5770, 5774, 5782: ~1.4 M words) are material for a spelling/normalization layer or decoder biasing, and for consistent abbreviation forms.
Compute. The large-v3 runs take 30+ h per Mac Studio; the GB10 (CUDA, bf16) would cut that to hours once its GPU memory and disk are freed. Fleet alignment should move from static shards to a shared queue (the tail idled the fleet ~1 h per wave).
Product and review. Synced subtitles (an SRT per file exists) are the preferred first deliverable; nothing is published or made searchable until Rabonim/Mashpi'im and the hanachos' author have reviewed; ask The Daily Sicha team for the 5782 recordings (text exists) and confirm the two 5781 days whose printed text does not match the audio (17 Kislev, 9 Teves).
Housekeeping. Report every number from the GPU path only (faster-whisper long-form loses ~3 points and suppresses the gershayim by default); keep output_superseded/ and the windowing reports as the audit trail.