Training-data sheet · snapshot 2026-09-24 07:20 EDT (v4 complete, v5 data building)Sichos / asr · prepared for the training scientist

Export: this page prints cleanly (File → Print → Save as PDF; the layout has print styles). A PDF, the standalone HTML, every chart as SVG and PNG, and the underlying tables as CSV are in asr/reports/corpus-datasheet-export.zip.

Rebbe Yiddish ASR Corpus

Everything the speech-recognition models in this project train and test on, updated 2026-09-24 with the new source data, the v4 results, the audit of the delivered transcripts and the plan for v5: where the audio and text come from, how they were paired and cut into clips, what was filtered, how the fixed evaluation sets were chosen, the pseudo-labeled extension, and what each training run did with it. All figures are computed from the manifests in asr/data/manifests/ and the evaluation reports in asr/reports/; nothing here is estimated.

1,468paired days (audio + Yiddish hanacha)
30,212clips ≤ 28 s, 168 h
1.25 Mreference words
27,833clean training clips, 153 h
20 / 20fixed dev / test days
+1,758paired days arriving in v5 (5775–5781): 306 h, 2.31 M words

Hours are audio inside clips. Words are counted on the training-target text (punctuation kept, vowel points removed). Hebrew years: 5711 = 1950/51 … 5752 = 1991/92; 5783 = 2022/23 … 5787 = 2026/27.

What one training example is

A clip is a stretch of at most 28 s of the Rebbe's speech, cut from a Daily Sicha recording, paired with the words the published Yiddish hanacha (the written transcript of the talk) gives for that stretch. The model sees the 16 kHz mono audio and learns to emit the text. A typical clip is 20 s long and carries about 42 words at 2.1 words per second.

id 5783/001/000 · 38.22–54.59 s · 16.4 s · timing: site sync (offset-corrected)

[the example text is withheld on the public page: the hanachos belong to The Daily Sicha and are not republished here]

One real training target. Note the conventions the model must learn: Yiddish in unpointed Hebrew script, loshon-kodesh phrases embedded verbatim, source names abbreviated with gershayim (רמב"ן, של"ה), quoted terms in quotation marks.

Where the data comes from

The Daily Sicha (thedailysicha.com) publishes one ~10-minute excerpt of a farbrengen per day with a Yiddish hanacha, and for recent years a sync file with segment timings. The project holds the daily mp3s for five publication years (5783–5787) and the site's text for each; the source talks span 1950 to 1992. Four older archive years (5778–5781, 1,005 files) have audio only and are the transcription target of the project.

Pair daysResolve each mp3 to its Hebrew date, fetch the day's hanacha and sync JSON.fetch_days.py
Measure offsets292 of 613 daily files carry a dedication before the excerpt (median 9.2 s); cross-correlation finds where the excerpt starts.compute_offsets.py
Align text to audioWhere the site has no timings, a Yiddish Whisper model transcribes with word times and the hypothesis is anchored to the reference (exact, then fuzzy).align_day.py
Cut clipsMerge consecutive timed segments up to 28 s; write FLAC clips and manifests; split by day.build_dataset.py
FilterDrop clips whose baseline transcription disagrees wildly with the reference (misaligned timings).filter_manifest.py
ExtendTranscribe the 1,005 text-less files with the best model and turn confident windows into pseudo-labeled clips.build_pseudo_dataset.py
0 10 20 30 40 hours of clips 5783 · hours: 26.65 26.65 5783 242 days · 4,873 clips 5784 · hours: 34.02 34.02 5784 318 days · 6,290 clips 5785 · hours: 33.23 33.23 5785 290 days · 6,080 clips 5786 · hours: 36.51 36.51 5786 291 days · 6,224 clips 5787 · hours: 37.38 37.38 5787 319 days · 6,745 clips
Clips by Daily Sicha publication year. 1,460 of the 1,468 paired days yielded clips (8 days failed alignment). Per year: 5783 = 242 days, 5784 = 318, 5785 = 290, 5786 = 291, 5787 = 319.
0 10,000 20,000 30,000 clips train · own anchor alignment: 27,461 train · site sync timings (offset-corrected): 1,812 29,273 train dev · own anchor alignment: 451 dev · site sync timings (offset-corrected): 0 451 dev test · own anchor alignment: 0 test · site sync timings (offset-corrected): 488 488 test
How each clip got its timings. 27,912 clips (92%) are timed by the project's own anchor alignment, 2,300 (8%) by the site's sync files corrected for the dedication offset. The test set is deliberately all site-timed; dev is all self-aligned (see splits).
0 2.5 5 7.5 10 12.5 hours of paired audio 5711 · hours: 1.18 5712 · hours: 1.12 5712 5713 · hours: 0.01 5714 · hours: 2.04 5715 · hours: 1.7 5715 5716 · hours: 2.76 5717 · hours: 1.89 5718 · hours: 2.21 5718 5719 · hours: 1.48 5720 · hours: 2.27 5721 · hours: 2.08 5721 5722 · hours: 3.17 5723 · hours: 2.05 5724 · hours: 2.13 5724 5725 · hours: 3.31 5726 · hours: 1.8 5727 · hours: 1.87 5727 5728 · hours: 2.5 5729 · hours: 2.45 5730 · hours: 1.56 5730 5731 · hours: 2.37 5732 · hours: 2.65 5733 · hours: 2.46 5733 5734 · hours: 3.72 5735 · hours: 4.8 5736 · hours: 6.99 5736 5737 · hours: 4.67 5738 · hours: 7.86 5739 · hours: 11.81 5739 5740 · hours: 8.99 5741 · hours: 7.57 5742 · hours: 6.45 5742 5743 · hours: 8.89 5744 · hours: 7.68 5745 · hours: 9.39 5745 5746 · hours: 8.53 5747 · hours: 6.03 5748 · hours: 3.76 5748 5749 · hours: 5.02 5750 · hours: 1.59 5751 · hours: 1.56 5751 5752 · hours: 1.86 5739 and later (the years the hanachos' author asked for)
Hours of paired audio by year of the original farbrengen (5711–5752; 637 clips from 32 days have no recorded year and are omitted here). The bulk of the material is 5734–5749; the shaded region marks 5739 onward, the years the requested transcriptions must cover. Older recordings are noisier and the speaker is younger, so the model sees both conditions.

New since 2026-09-23: hanachos and audio for 5775–5781

The Daily Sicha delivered the Yiddish hanachos as one PDF per year (5768–5782 and 5784–5787) together with the recordings of 5775–5777. The four years the project had been transcribing without text (5778–5781) and three more years are now paired: 1,758 days, 306 h of recordings, 2.31 M words, more than doubling the corpus once aligned. Nine PDFs have a clean text layer. Ten (5768–5770, 5774–5780) are set in custom-encoded fonts: every glyph code is a fixed permutation of the Hebrew alphabet, different per font. They are decoded exactly rather than OCR'd: pipeline/pdf_fontmap.py hill-climbs each font's code→letter map by the share of decoded words that exist in the training vocabulary (0.91–0.93 for the body font) and pins the bold header font with the words every day's header shares (בס"ד. "השיחה היומית" ליום, the year, התוכן, the month names). Tesseract on the same pages scored 0.39 WER; the exact decode leaves no systematic errors (the drafts-vs-text numbers below equal the clean-layer year).

Decode the fontsLearn one code→letter map per font from ~100k sample words; bold font pinned by the header crib.pdf_fontmap.py
Split into daysHeader line → day; the printed [NNN] number → the mp3 stem. All 1,758 days matched, none missing or duplicated.pdf_hanachos.py
Day recordsSame record shape as the site years, text as one paragraph per PDF line, no site timings (offset 0).build_pdf_days.py
Align on the fleet7 Mac Studios × 5 workers, production CT2 model anchors the text to the audio (anchor rates 0.94–0.98).align_day.py
Build + scanClips ≤ 28 s; the production model transcribes every train clip; clips with CER > 0.6 are dropped (per-year drop rate = text-quality signal).build_dataset.py · filter_manifest.py
new years (aligning now)v4 training years
0 100 200 300 400 paired days per Daily Sicha publication year 5775 · 42.5 h audio · 0.32 M words · decoded : 247.000 247.000 5776 · 46.5 h audio · 0.35 M words · decoded : 266.000 266.000 5777 · 42.3 h audio · 0.32 M words · decoded : 240.000 240.000 5778 · 43.4 h audio · 0.33 M words · decoded : 247.000 247.000 5779 · 46.6 h audio · 0.35 M words · decoded : 266.000 266.000 5780 · 42.1 h audio · 0.32 M words · decoded : 245.000 245.000 5781 · 43.1 h audio · 0.33 M words · clean text layer : 247.000 247.000 5783 · site text (v4 training data) : 242.000 242.000 5784 · site text (v4 training data) : 318.000 318.000 5785 · site text (v4 training data) : 290.000 290.000 5786 · site text (v4 training data) : 291.000 291.000 5787 · site text (v4 training data) : 319.000 319.000
Paired days per Daily Sicha publication year. Hours are recording time before clipping (roughly 65% of it ends up inside clips). Text for 5768–5770, 5774 and 5782 exists too (about 1.4 M words) but no audio is in hand, so those years are text-only for now.

Three traps in the PDFs, all fixed at the glyph level. (1) 5776–5778 carry a diagonal watermark (הנחה פרטית בלתי מוגה, 80–85 pt): its letters landed on body rows and fused words (2,100–2,700 fourteen-letter "words" per year against ~130 in clean years) and left ~500 stray single letters per year; glyphs over 40 pt are now dropped before rows are formed. (2) Every page's footer small print (contact line, 6–10 pt, Arial/Tahoma) decoded through the wrong font into junk tokens, up to 2,788 per year; rows are now kept only in the body font family, not by size (two 10 pt days would have been emptied by a size cut), and the training-text normalizer drops any token with Latin letters. (3) The bold header font learned without the crib landed in a wrong permutation on 5775/5776 (0.03 in-vocabulary, zero days found); with the crib it reaches 0.82–0.85. After the fixes the 5778 drafts score 0.072 against the text instead of 0.120: the gap was the text, not the audio.

Splits

Splitting is by day, never by clip, so no test audio shares a recording with training audio. The rule, fixed on 2026-09-16 and never changed:

test: 20 evenly spaced days among the site-timed days (Tamuz 5786-Tishrei 5787) whose source farbrengen is 5739 or later and whose offset is known; dev: 20 evenly spaced days among the remaining text-only days with known offsets. Fixed on 2026-09-16; never train on these days.

splitdaysclipshoursreference wordstiming source
train (all)1,35229,273162.01,210,72527,461 self-aligned · 1,812 site-timed
train.clean (used by every v1–v3 run)1,35227,833≈153≈1,161,000after the CER filter below
dev204512.4118,371self-aligned
test204883.3423,547site-timed, farbrengens 5739+

Dev days: 001 3-Tishrei, 028 11-Cheshvan, 055 13-Kislev, 082 15-Teiveis, 109 17-Shvat, 137 20-Adar, 164 27-Nisan, 191 28-Iyar, 023 2-Cheshvan 5787, 050 3-Kisleiv 5787, 077 5-Teiveis 5787, 104 7-Shvat 5787, 131 9-Adar 1 5787, 158 10-Adar 2 5787, 185 13-Nisan 5787, 213 20-Iyar 5787, 240 24-Sivan 5787, 267 25-Tamuz 5787, 294 28-Av 5787, 321 29-Elul 5787

Test days: 227 13-Tamuz, 237 24-Tamuz, 242 1-Av, 246 6-Av, 251 12-Av, 254 15-Av, 260 22-Av, 263 26-Av, 265 28-Av, 268 1-Elul, 272 6-Elul, 278 13-Elul, 281 17-Elul, 284 20-Elul, 001 3-Tishrei 5787, 006 9-Tishrei 5787, 009 13-Tishrei 5787, 016 24-Tishrei 5787, 018 26-Tishrei 5787, 021 30-Tishrei 5787

Caveats for the scientist: (1) the dev WER printed during training uses a fixed 200-clip subset of dev decoded with the HF generate path; the full-dev and test numbers in the results section use the full sets. (2) One test day, 9 Tishrei 5787, scores 0.28–0.34 for every model and is a suspected reference-quality outlier; it is kept because the split is frozen. (3) Out-of-vocabulary rate against the clean training text: dev 1.1%, test 1.2%.

Clip anatomy

0 2,000 4,000 6,000 0–2: 43 clips 2–4: 210 clips 4–6: 432 clips 6–8: 660 clips 8–10: 946 clips 10–12: 1,327 clips 12–14: 1,763 clips 14–16: 2,167 clips 16–18: 2,495 clips 18–20: 2,992 clips 20–22: 3,385 clips 22–24: 3,967 clips 24–26: 4,744 clips 26–28: 5,081 clips 0 2 4 6 8 10 12 14 16 18 20 22 24 26 28 clip duration (s)clips
Clip duration. Mean 20.0 s, median 21.2 s, 10th–90th percentile 11.1–26.8 s, max 28.3 s. The builder merges consecutive timed segments until the next one would exceed 28 s, so most clips sit just under the Whisper 30 s window.
0 2,000 4,000 6,000 0–5: 97 clips 5–10: 365 clips 10–15: 701 clips 15–20: 1,110 clips 20–25: 1,595 clips 25–30: 2,191 clips 30–35: 2,880 clips 35–40: 3,689 clips 40–45: 4,200 clips 45–50: 4,378 clips 50–55: 3,779 clips 55–60: 2,605 clips 60–65: 1,445 clips 65–70: 670 clips 70–75: 322 clips 75–80: 107 clips 80–85: 40 clips 85–90: 16 clips 90–95: 11 clips 95–100: 5 clips 0 10 20 30 40 50 60 70 80 90 100 words per clipclips
Words per clip. Mean 41.5, median 43; 2.07 words per second of audio overall. Clips per day: median 21 (1–57).

Alignment quality

For the 92% of clips timed by the project's own aligner, the anchor rate is the share of a clip's reference words that were matched exactly or fuzzily to a word the Yiddish Whisper model heard at a known time; the remaining words are placed by interpolation. Clips below 0.30 are excluded at build time, and clip boundaries must fall on anchored words.

0 2,000 4,000 6,000 0.3–0.35: 306 clips 0.35–0.4: 500 clips 0.4–0.45: 929 clips 0.45–0.5: 1,244 clips 0.5–0.55: 2,732 clips 0.55–0.6: 3,390 clips 0.6–0.65: 4,268 clips 0.65–0.7: 4,393 clips 0.7–0.75: 3,796 clips 0.75–0.8: 3,099 clips 0.8–0.85: 1,934 clips 0.85–0.9: 829 clips 0.9–0.95: 325 clips 0.95–1: 167 clips 0.3 0.4 0.5 0.6 0.7 0.8 0.9 1 anchor rate of the clip's wordsclips
Anchor rate of self-aligned clips (n = 27,912; mean 0.65, median 0.66). Validation against the site's own timings on 3 Tishrei 5787 (113 segments): median start difference 0.33 s, 90th percentile 0.95 s, 83% of segments within 1 s at both ends.

Cleaning

The whole train split was transcribed once with the untouched Yiddish model and each clip's hypothesis compared with its reference. Clips with character error rate above 0.60, or a hypothesis/reference length ratio outside 0.5–2.0, were removed: 1,440 clips (4.9%), 1,436 of them by the CER rule (median CER of the dropped clips 0.75). These are almost always timing errors, where the text does not belong to the audio, not hard speech. The v0 model trained on the unfiltered set; every run since uses train.clean.jsonl.

Vocabulary

Clean training text: 1,161,335 tokens, 27,712 distinct words (letters only, no marks), 10,646 of them seen once. The 1,000 most frequent words cover 81.8% of tokens, the top 5,000 cover 94.3%. Words seen 100+ times are 4.0% of the types but 82.7% of the tokens — the long tail of names and sources is where the remaining errors live.

Text conventions and normalization

Training target
normalize.training_text: the hanacha paragraph text with vowel points (nikkud) removed, punctuation and the abbreviation marks kept (geresh ' and gershayim "), bracket characters removed but the bracketed words kept (they are spoken asides; dropping them raised WER for every model). 27,644 of 30,212 raw references carried nikkud; 68.8% of clips contain at least one gershayim abbreviation.
Scoring text
normalize.scoring_text: Hebrew letters only. All WER and CER figures anywhere in the project are computed on this form, so punctuation, quotes and spelling of marks never count as errors, but a misspelled abbreviation does.
Reference frames
Site sync timings can refer to the raw excerpt or to the daily file with its dedication; calibrate_site_timings.py decides per day and compute_offsets.py supplies the offset. Days whose offset could not be established are never used with site timings.
Decoder note
Whisper's default token-suppression list blocks the ASCII double quote, which in Hebrew text is the gershayim. Every fine-tuned model learned to write רש"י and הקב"ה but could not emit them through faster-whisper until 2026-09-21; the production decoder (HF fp32 on the GPU, VAD windows, no timestamp tokens) has no such list.

Pseudo-labeled extension (v3 self-training)

Superseded for training on 2026-09-23: the 5778–5781 recordings now have their real hanachos (section above), so v5 trains on human text for these years; the pseudo-labels remain as a documented experiment. Result of self-training (v3 + v4): it lifts weaker starting points by 0.5–1 point but never beats the best supervised recipe on the same data (production 0.081 → 0.081; large lr 2e-5 0.082 → 0.085).

A second pseudo-label set (26,797 clips, 169.4 h) was cut on 2026-09-22 from the third transcription pass (model turbo lr 3e-5, whole-file WER 0.081); the two v4 self-training runs use it, the v3 runs below used the first set. Result of v3 on the full test set: 0.082 / 0.084 / 0.086 whole-file for the three recipes — each 0.5–1 point better than its starting model, none better than the supervised lr 3e-5 model (0.081).

The 1,005 recordings of 5778–5781 have no human text. They were transcribed with the best model (turbo full fine-tune, lr 2e-5) through the production decoder; each VAD window of up to 28 s became a clip with the model's text as its label. Windows with mean token log-probability below −0.6, fewer than two words, or repetitive text (decoder loops, 32 windows) were dropped, and labels were normalized exactly like human targets.

0 10 20 30 40 50 hours of pseudo-labeled clips DS 5778 · hours: 42 42 DS 5778 247 files · 6,687 clips Sicha Yomis 5779 · hours: 44.9 44.9 Sicha Yomis 5779 266 files · 7,089 clips Sicha Yomis 5780 · hours: 40.7 40.7 Sicha Yomis 5780 245 files · 6,446 clips Sicha Yomis 5781 · hours: 41.6 41.6 Sicha Yomis 5781 247 files · 6,556 clips
Pseudo-labeled clips by archive folder. 26,778 clips, 169.3 h, mean 22.8 s, 77,201 gershayim in the labels. Confidence (mean token log-prob per window): 10th percentile −0.065, median −0.035, 90th −0.020.

Mixed training set for v3

train.clean_pseudo.jsonl = 27,833 human-timed clips + 26,778 pseudo-labeled clips = 54,611 rows, about 322 h. At batch 8 × 4 accumulation that is 1,707 steps per epoch. The pseudo-labels inherit the labeling model's error rate (about 9% of words), so v3 tests whether in-domain audio with imperfect labels beats the same continuation on real data alone (run on bigmac03 in v2 for exactly this comparison).

Rows carry timing_source: pseudo, min_logprob and n_segments, so they can be filtered or down-weighted later without rebuilding.

Training runs

All runs fine-tune ivrit.ai's Yiddish-adapted Whisper (large-v3 or large-v3-turbo) in fp32 on Apple-silicon GPUs, one run per Mac Studio, from the same clean manifest; batch 8 × 4 accumulation (32 clips per step), 870 steps per epoch, linear decay with 200 warm-up steps, evaluation every 500 steps on the 200-clip dev subset. Lines are dev WER; a dot is an evaluation.

v1 · 8 recipes, complete (2026-09-19 to 21) 0.06 0.11 0.16 0.22 0.27 0.32 0 1000 2000 3000 optimizer step turbo · full FT · lr 2e-5 · 3 ep turbo · full FT · lr 2e-5 · 3 ep step 500: 0.1261 turbo · full FT · lr 2e-5 · 3 ep step 1000: 0.1154 turbo · full FT · lr 2e-5 · 3 ep step 1500: 0.0918 turbo · full FT · lr 2e-5 · 3 ep step 2000: 0.0808 turbo · full FT · lr 2e-5 · 3 ep step 2500: 0.0820 turbo · full FT · lr 2e-5 · 3 ep step 2610: 0.0765 turbo · full FT · lr 1e-5 · 3 ep turbo · full FT · lr 1e-5 · 3 ep step 500: 0.1258 turbo · full FT · lr 1e-5 · 3 ep step 1000: 0.1133 turbo · full FT · lr 1e-5 · 3 ep step 1500: 0.1026 turbo · full FT · lr 1e-5 · 3 ep step 2000: 0.0878 turbo · full FT · lr 1e-5 · 3 ep step 2500: 0.0804 turbo · full FT · lr 1e-5 · 3 ep step 2610: 0.0793 turbo · LoRA r64 · 3 ep turbo · LoRA r64 · 3 ep step 500: 0.1346 turbo · LoRA r64 · 3 ep step 1000: 0.1105 turbo · LoRA r64 · 3 ep step 1500: 0.0985 turbo · LoRA r64 · 3 ep step 2000: 0.0885 turbo · LoRA r64 · 3 ep step 2500: 0.0843 turbo · LoRA r64 · 3 ep step 2610: 0.0836 turbo · frozen encoder · lr 2e-5 · 4 ep turbo · frozen encoder · lr 2e-5 · 4 ep step 500: 0.2806 turbo · frozen encoder · lr 2e-5 · 4 ep step 1000: 0.2679 turbo · frozen encoder · lr 2e-5 · 4 ep step 1500: 0.2283 turbo · frozen encoder · lr 2e-5 · 4 ep step 2000: 0.2218 turbo · frozen encoder · lr 2e-5 · 4 ep step 2500: 0.1945 turbo · frozen encoder · lr 2e-5 · 4 ep step 3000: 0.2160 turbo · frozen encoder · lr 2e-5 · 4 ep step 3480: 0.1959 turbo · frozen encoder · lr 1e-5 · 4 ep turbo · frozen encoder · lr 1e-5 · 4 ep step 500: 0.3122 turbo · frozen encoder · lr 1e-5 · 4 ep step 1000: 0.2751 turbo · frozen encoder · lr 1e-5 · 4 ep step 1500: 0.2725 turbo · frozen encoder · lr 1e-5 · 4 ep step 2000: 0.2311 turbo · frozen encoder · lr 1e-5 · 4 ep step 2500: 0.2220 turbo · frozen encoder · lr 1e-5 · 4 ep step 3000: 0.2341 turbo · frozen encoder · lr 1e-5 · 4 ep step 3480: 0.2277 large-v3 · full FT · lr 1e-5 · 1 ep large-v3 · full FT · lr 1e-5 · 1 ep step 500: 0.1179 large-v3 · full FT · lr 1e-5 · 1 ep step 870: 0.1025 large-v3 · LoRA r64 · 1 ep large-v3 · LoRA r64 · 1 ep step 500: 0.1254 large-v3 · LoRA r64 · 1 ep step 870: 0.1029 large-v3 · frozen encoder · 2 ep large-v3 · frozen encoder · 2 ep step 500: 0.1807 large-v3 · frozen encoder · 2 ep step 1000: 0.1689 large-v3 · frozen encoder · 2 ep step 1500: 0.1664 large-v3 · frozen encoder · 2 ep step 1740: 0.1664
  • turbo · full FT · lr 2e-5 · 3 ep 2610 steps · last dev 0.0765
  • turbo · full FT · lr 1e-5 · 3 ep 2610 steps · last dev 0.0793
  • turbo · LoRA r64 · 3 ep 2610 steps · last dev 0.0836
  • turbo · frozen encoder · lr 2e-5 · 4 ep 3480 steps · last dev 0.1959
  • turbo · frozen encoder · lr 1e-5 · 4 ep 3480 steps · last dev 0.2277
  • large-v3 · full FT · lr 1e-5 · 1 ep 870 steps · last dev 0.1025
  • large-v3 · LoRA r64 · 1 ep 870 steps · last dev 0.1029
  • large-v3 · frozen encoder · 2 ep 1740 steps · last dev 0.1664
v2 · 4 recipes, complete (2026-09-21 to 22) 0.06 0.08 0.10 0.12 0.14 0.16 0 1000 2000 3000 optimizer step best v1 final 0.0765 continue best v1 (turbo lr 2e-5) · lr 1e-5 · 3 ep continue best v1 (turbo lr 2e-5) · lr 1e-5 · 3 ep step 500: 0.0844 continue best v1 (turbo lr 2e-5) · lr 1e-5 · 3 ep step 1000: 0.0905 continue best v1 (turbo lr 2e-5) · lr 1e-5 · 3 ep step 1500: 0.0784 continue best v1 (turbo lr 2e-5) · lr 1e-5 · 3 ep step 2000: 0.0766 continue best v1 (turbo lr 2e-5) · lr 1e-5 · 3 ep step 2500: 0.0757 continue best v1 (turbo lr 2e-5) · lr 1e-5 · 3 ep step 2610: 0.0805 turbo · full FT · lr 3e-5 · 4 ep turbo · full FT · lr 3e-5 · 4 ep step 500: 0.1396 turbo · full FT · lr 3e-5 · 4 ep step 1000: 0.1334 turbo · full FT · lr 3e-5 · 4 ep step 1500: 0.0966 turbo · full FT · lr 3e-5 · 4 ep step 2000: 0.0884 turbo · full FT · lr 3e-5 · 4 ep step 2500: 0.0777 turbo · full FT · lr 3e-5 · 4 ep step 3000: 0.0746 turbo · full FT · lr 3e-5 · 4 ep step 3480: 0.0730 large-v3 · full FT · lr 1e-5 · 3 ep large-v3 · full FT · lr 1e-5 · 3 ep step 500: 0.1163 large-v3 · full FT · lr 1e-5 · 3 ep step 1000: 0.0961 large-v3 · full FT · lr 1e-5 · 3 ep step 1500: 0.0911 large-v3 · full FT · lr 1e-5 · 3 ep step 2000: 0.0826 large-v3 · full FT · lr 1e-5 · 3 ep step 2500: 0.0816 large-v3 · full FT · lr 1e-5 · 3 ep step 2610: 0.0809 large-v3 · full FT · lr 2e-5 · 3 ep large-v3 · full FT · lr 2e-5 · 3 ep step 500: 0.1266 large-v3 · full FT · lr 2e-5 · 3 ep step 1000: 0.0878 large-v3 · full FT · lr 2e-5 · 3 ep step 1500: 0.0843 large-v3 · full FT · lr 2e-5 · 3 ep step 2000: 0.0776 large-v3 · full FT · lr 2e-5 · 3 ep step 2500: 0.0771 large-v3 · full FT · lr 2e-5 · 3 ep step 2610: 0.0723
  • continue best v1 (turbo lr 2e-5) · lr 1e-5 · 3 ep 2610 steps · last dev 0.0805
  • turbo · full FT · lr 3e-5 · 4 ep 3480 steps · last dev 0.0730
  • large-v3 · full FT · lr 1e-5 · 3 ep 2610 steps · last dev 0.0809
  • large-v3 · full FT · lr 2e-5 · 3 ep 2610 steps · last dev 0.0723
v3 · self-training, complete (2026-09-21 to 22) 0.06 0.08 0.09 0.11 0.12 0.14 0 1000 2000 3000 optimizer step best v1 final 0.0765 continue v1 lr 2e-5 on real + pseudo · lr 1e-5 · 2 ep continue v1 lr 2e-5 on real + pseudo · lr 1e-5 · 2 ep step 500: 0.0803 continue v1 lr 2e-5 on real + pseudo · lr 1e-5 · 2 ep step 1000: 0.0760 continue v1 lr 2e-5 on real + pseudo · lr 1e-5 · 2 ep step 1500: 0.0768 continue v1 lr 2e-5 on real + pseudo · lr 1e-5 · 2 ep step 2000: 0.0708 continue v1 lr 2e-5 on real + pseudo · lr 1e-5 · 2 ep step 2500: 0.0766 continue v1 lr 2e-5 on real + pseudo · lr 1e-5 · 2 ep step 3000: 0.0686 continue v1 lr 2e-5 on real + pseudo · lr 1e-5 · 2 ep step 3414: 0.0708 continue v1 lr 1e-5 on real + pseudo · lr 1e-5 · 2 ep continue v1 lr 1e-5 on real + pseudo · lr 1e-5 · 2 ep step 500: 0.0892 continue v1 lr 1e-5 on real + pseudo · lr 1e-5 · 2 ep step 1000: 0.0946 continue v1 lr 1e-5 on real + pseudo · lr 1e-5 · 2 ep step 1500: 0.0857 continue v1 lr 1e-5 on real + pseudo · lr 1e-5 · 2 ep step 2000: 0.0763 continue v1 lr 1e-5 on real + pseudo · lr 1e-5 · 2 ep step 2500: 0.0752 continue v1 lr 1e-5 on real + pseudo · lr 1e-5 · 2 ep step 3000: 0.0709 continue v1 lr 1e-5 on real + pseudo · lr 1e-5 · 2 ep step 3414: 0.0693 fresh turbo on real + pseudo · lr 2e-5 · 2 ep fresh turbo on real + pseudo · lr 2e-5 · 2 ep step 500: 0.1330 fresh turbo on real + pseudo · lr 2e-5 · 2 ep step 1000: 0.1019 fresh turbo on real + pseudo · lr 2e-5 · 2 ep step 1500: 0.1146 fresh turbo on real + pseudo · lr 2e-5 · 2 ep step 2000: 0.0778 fresh turbo on real + pseudo · lr 2e-5 · 2 ep step 2500: 0.0762 fresh turbo on real + pseudo · lr 2e-5 · 2 ep step 3000: 0.0713 fresh turbo on real + pseudo · lr 2e-5 · 2 ep step 3414: 0.0686
  • continue v1 lr 2e-5 on real + pseudo · lr 1e-5 · 2 ep 3414 steps · last dev 0.0708
  • continue v1 lr 1e-5 on real + pseudo · lr 1e-5 · 2 ep 3414 steps · last dev 0.0693
  • fresh turbo on real + pseudo · lr 2e-5 · 2 ep 3414 steps · last dev 0.0686
v4 · 6 recipes, complete (2026-09-22 to 24) 0.06 0.09 0.12 0.14 0.17 0.20 0 1000 2000 3000 optimizer step best v1 final 0.0765 v4: production model self-trained on pass-3 labels · lr 1e-5 · 2 ep v4: production model self-trained on pass-3 labels · lr 1e-5 · 2 ep step 500: 0.0883 v4: production model self-trained on pass-3 labels · lr 1e-5 · 2 ep step 1000: 0.0778 v4: production model self-trained on pass-3 labels · lr 1e-5 · 2 ep step 1500: 0.0821 v4: production model self-trained on pass-3 labels · lr 1e-5 · 2 ep step 2000: 0.0740 v4: production model self-trained on pass-3 labels · lr 1e-5 · 2 ep step 2500: 0.0731 v4: production model self-trained on pass-3 labels · lr 1e-5 · 2 ep step 3000: 0.0713 v4: production model self-trained on pass-3 labels · lr 1e-5 · 2 ep step 3416: 0.0699 v4: large lr 2e-5 + pass-3 pseudo · lr 1e-5 · 2 ep v4: large lr 2e-5 + pass-3 pseudo · lr 1e-5 · 2 ep step 500: 0.0746 v4: large lr 2e-5 + pass-3 pseudo · lr 1e-5 · 2 ep step 1000: 0.0707 v4: large lr 2e-5 + pass-3 pseudo · lr 1e-5 · 2 ep step 1500: 0.0869 v4: large lr 2e-5 + pass-3 pseudo · lr 1e-5 · 2 ep step 2000: 0.0707 v4: large lr 2e-5 + pass-3 pseudo · lr 1e-5 · 2 ep step 2500: 0.0727 v4: large lr 2e-5 + pass-3 pseudo · lr 1e-5 · 2 ep step 3000: 0.0713 v4: large lr 2e-5 + pass-3 pseudo · lr 1e-5 · 2 ep step 3416: 0.0733 turbo · full FT · lr 4e-5 · 4 ep turbo · full FT · lr 4e-5 · 4 ep step 500: 0.1598 turbo · full FT · lr 4e-5 · 4 ep step 1000: 0.1125 turbo · full FT · lr 4e-5 · 4 ep step 1500: 0.1002 turbo · full FT · lr 4e-5 · 4 ep step 2000: 0.1015 turbo · full FT · lr 4e-5 · 4 ep step 2500: 0.0840 turbo · full FT · lr 4e-5 · 4 ep step 3000: 0.0820 turbo · full FT · lr 4e-5 · 4 ep step 3480: 0.0847 turbo · full FT · lr 5e-5 · 4 ep turbo · full FT · lr 5e-5 · 4 ep step 500: 0.1803 turbo · full FT · lr 5e-5 · 4 ep step 1000: 0.1244 turbo · full FT · lr 5e-5 · 4 ep step 1500: 0.1079 turbo · full FT · lr 5e-5 · 4 ep step 2000: 0.0943 turbo · full FT · lr 5e-5 · 4 ep step 2500: 0.0889 turbo · full FT · lr 5e-5 · 4 ep step 3000: 0.0765 turbo · full FT · lr 5e-5 · 4 ep step 3480: 0.0713 large-v3 · full FT · lr 3e-5 · 4 ep large-v3 · full FT · lr 3e-5 · 4 ep step 500: 0.1171 large-v3 · full FT · lr 3e-5 · 4 ep step 1000: 0.0971 large-v3 · full FT · lr 3e-5 · 4 ep step 1500: 0.0846 large-v3 · full FT · lr 3e-5 · 4 ep step 2000: 0.0803 large-v3 · full FT · lr 3e-5 · 4 ep step 2500: 0.0782 large-v3 · full FT · lr 3e-5 · 4 ep step 3000: 0.0738 large-v3 · full FT · lr 3e-5 · 4 ep step 3480: 0.0751 turbo · full FT · lr 3e-5 · 4 ep · seed 7 (repeat) turbo · full FT · lr 3e-5 · 4 ep · seed 7 (repeat) step 500: 0.1383 turbo · full FT · lr 3e-5 · 4 ep · seed 7 (repeat) step 1000: 0.1111 turbo · full FT · lr 3e-5 · 4 ep · seed 7 (repeat) step 1500: 0.0944 turbo · full FT · lr 3e-5 · 4 ep · seed 7 (repeat) step 2000: 0.0907 turbo · full FT · lr 3e-5 · 4 ep · seed 7 (repeat) step 2500: 0.0813 turbo · full FT · lr 3e-5 · 4 ep · seed 7 (repeat) step 3000: 0.0782 turbo · full FT · lr 3e-5 · 4 ep · seed 7 (repeat) step 3480: 0.0740
  • v4: production model self-trained on pass-3 labels · lr 1e-5 · 2 ep 3416 steps · last dev 0.0699
  • v4: large lr 2e-5 + pass-3 pseudo · lr 1e-5 · 2 ep 3416 steps · last dev 0.0733
  • turbo · full FT · lr 4e-5 · 4 ep 3480 steps · last dev 0.0847
  • turbo · full FT · lr 5e-5 · 4 ep 3480 steps · last dev 0.0713
  • large-v3 · full FT · lr 3e-5 · 4 ep 3480 steps · last dev 0.0751
  • turbo · full FT · lr 3e-5 · 4 ep · seed 7 (repeat) 3480 steps · last dev 0.0740

Reading the curves: freezing the encoder (amber) plateaus at 0.19–0.22, LoRA (teal, purple) lands near 0.084, full fine-tuning of the turbo model reaches 0.0765 at lr 2e-5. v2 finished with large-v3 at lr 2e-5 on 0.0723 and turbo at lr 3e-5 on 0.0730 (both better than the v1 winner); continuing the v1 winner at lr 1e-5 only reached 0.0757. The v3 self-training runs ended with the best in-training scores of all (0.0686–0.0693) but scored 0.082–0.086 whole-file on the full test set, behind the production model — the 200-clip dev subset overstates models trained on pseudo-labels, so only the full test set ranks them. Pseudo-labels lifted each starting model by 0.5–1 point without surpassing the best supervised recipe. v4 probes higher learning rates, the turbo recipe on large-v3, a seed repeat of the production run, and self-training from the strongest starting points with fresher labels.

Results on the fixed test set

Two metrics, both against the human hanacha of the 20 test days: clip WER on the 488 test clips (audio cut on known timings) and whole-file WER where each day's full recording is decoded with voice-activity chunking and scored against the whole hanacha — the number that matters for the product. Decoder matters as much as model: the same weights score 0.132, 0.124 and 0.088 whole-file under the three decoders tried. The v2 winner on whole files is the turbo model trained at lr 3e-5 for 4 epochs (0.081, CER 0.045); large-v3 at lr 2e-5 is best on clips (0.083) but decodes three times slower and is a hair worse on whole files (0.082).

clip test WERwhole-file test WER
0 0.1 0.2 0.3 0.4 0.5 0.6 word error rate on the 20 fixed test days (lower is better) v2: large-v3 full lr 2e-5, 3 ep · GPU path clip test WER: 0.083 0.083 whole-file test WER: 0.082 0.082 v4: large-v3 full lr 3e-5, 4 ep (resumed after the disk crash) · GPU path clip test WER: 0.086 0.086 whole-file test WER: 0.081 0.081 v2: turbo full lr 3e-5, 4 ep · GPU path (production since 09-22) clip test WER: 0.087 0.087 whole-file test WER: 0.081 0.081 v4: production model self-trained on pass-3 labels · GPU path clip test WER: 0.087 0.087 whole-file test WER: 0.081 0.081 v4: production recipe, seed 7 (repeat) · GPU path clip test WER: 0.087 0.087 whole-file test WER: 0.080 0.080 v4: turbo full lr 5e-5, 4 ep · GPU path clip test WER: 0.089 0.089 whole-file test WER: 0.082 0.082 v4: large-v3 lr 2e-5 model + pass-3 pseudo-labels · GPU path clip test WER: 0.089 0.089 whole-file test WER: 0.085 0.085 v3: turbo lr 1e-5 model + 2 ep on real + pseudo · GPU path clip test WER: 0.090 0.090 whole-file test WER: 0.082 0.082 v4: turbo full lr 4e-5, 4 ep · GPU path clip test WER: 0.090 0.090 whole-file test WER: 0.088 0.088 v3: turbo lr 2e-5 model + 2 ep on real + pseudo · GPU path clip test WER: 0.090 0.090 whole-file test WER: 0.084 0.084 v2: continue turbo lr 2e-5 at lr 1e-5, 3 ep · GPU path clip test WER: 0.091 0.091 whole-file test WER: 0.086 0.086 v2: large-v3 full lr 1e-5, 3 ep · GPU path clip test WER: 0.091 0.091 whole-file test WER: 0.087 0.087 v3: fresh turbo lr 2e-5, 2 ep on real + pseudo · GPU path clip test WER: 0.092 0.092 whole-file test WER: 0.086 0.086 turbo full lr 2e-5 · GPU path (production 09-21) clip test WER: 0.093 0.093 whole-file test WER: 0.088 0.088 turbo full lr 2e-5 · CT2 int8, quotes freed clip test WER: 0.095 0.095 whole-file test WER: 0.124 0.124 turbo full lr 1e-5 · GPU path clip test WER: 0.098 0.098 whole-file test WER: 0.093 0.093 turbo full lr 1e-5 · CT2 int8, quotes freed clip test WER: 0.101 0.101 whole-file test WER: 0.121 0.121 turbo full lr 2e-5 · CT2 int8, default suppression clip test WER: 0.105 0.105 whole-file test WER: 0.132 0.132 turbo full lr 1e-5 · CT2 default clip test WER: 0.110 0.110 whole-file test WER: 0.128 0.128 turbo LoRA r64 · CT2 default clip test WER: 0.118 0.118 whole-file test WER: 0.127 0.127 large-v3 full lr 1e-5, 1 ep · CT2 default clip test WER: 0.126 0.126 whole-file test WER: 0.144 0.144 large-v3 LoRA r64, 1 ep · CT2 default clip test WER: 0.134 0.134 whole-file test WER: 0.147 0.147 large-v3 frozen encoder · CT2 default clip test WER: 0.182 0.182 whole-file test WER: 0.199 0.199 turbo frozen encoder lr 2e-5 · CT2 default clip test WER: 0.225 0.225 whole-file test WER: 0.239 0.239 turbo frozen encoder lr 1e-5 · CT2 default clip test WER: 0.244 0.244 whole-file test WER: 0.266 0.266 v0: frozen encoder, 1,000 steps, unfiltered data clip test WER: 0.284 0.284 whole-file test WER: 0.305 0.305 baseline ivrit-ai yi-whisper-large-v3 (no fine-tuning) clip test WER: 0.531 0.531 whole-file test WER: 0.534 0.534 baseline ivrit-ai yi-whisper-large-v3-turbo (no fine-tuning) clip test WER: 0.571 0.571 whole-file test WER: 0.569 0.569 production model · auto windows (decoder default since 09-23) clip test WER: 0.000 0.000 whole-file test WER: 0.081 0.081
Baselines are the ivrit.ai models without any fine-tuning on this corpus. "CT2 int8" = CTranslate2/faster-whisper on CPU; "quotes freed" = same with the gershayim token un-suppressed; "GPU path" = HF fp32 model on MPS with VAD windows and no timestamp tokens (pipeline/hf_longform.py). Acceptance targets were ≤ 0.30 clip and ≤ 0.35 whole-file.
model · decoderclip dev WERclip test WERclip test CERwhole-file WERwhole-file CER
v2: turbo full lr 3e-5, 4 ep · GPU path (production since 09-22)0.0800.0870.0500.0810.045
production model · auto windows (decoder default since 09-23)–––0.0810.045
v4: large-v3 full lr 3e-5, 4 ep (resumed after the disk crash) · GPU path0.0800.0860.0510.0810.047
v4: production recipe, seed 7 (repeat) · GPU path0.0810.0870.0500.0800.045
v4: production model self-trained on pass-3 labels · GPU path0.0790.0870.0500.0810.045
v4: turbo full lr 5e-5, 4 ep · GPU path0.0790.0890.0530.0820.046
v4: turbo full lr 4e-5, 4 ep · GPU path0.0840.0900.0500.0880.048
v4: large-v3 lr 2e-5 model + pass-3 pseudo-labels · GPU path0.0800.0890.0500.0850.047
v2: large-v3 full lr 2e-5, 3 ep · GPU path0.0770.0830.0480.0820.047
v3: turbo lr 1e-5 model + 2 ep on real + pseudo · GPU path0.0790.0900.0510.0820.046
v3: turbo lr 2e-5 model + 2 ep on real + pseudo · GPU path0.0800.0900.0520.0840.046
v3: fresh turbo lr 2e-5, 2 ep on real + pseudo · GPU path0.0790.0920.0530.0860.047
v2: continue turbo lr 2e-5 at lr 1e-5, 3 ep · GPU path0.0820.0910.0510.0860.048
v2: large-v3 full lr 1e-5, 3 ep · GPU path0.0830.0910.0510.0870.048
turbo full lr 2e-5 · GPU path (production 09-21)0.0800.0930.0520.0880.048
turbo full lr 1e-5 · GPU path0.0880.0980.0530.0930.050
turbo full lr 2e-5 · CT2 int8, quotes freed0.0820.0950.0530.1240.081
turbo full lr 1e-5 · CT2 int8, quotes freed0.0930.1010.0570.1210.075
turbo full lr 2e-5 · CT2 int8, default suppression0.0920.1050.0570.1320.084
turbo full lr 1e-5 · CT2 default0.1000.1100.0600.1280.076
turbo LoRA r64 · CT2 default0.1010.1180.0640.1270.073
large-v3 full lr 1e-5, 1 ep · CT2 default0.1100.1260.0680.1440.083
large-v3 LoRA r64, 1 ep · CT2 default0.1150.1340.0710.1470.084
large-v3 frozen encoder · CT2 default0.1570.1820.0990.1990.117
turbo frozen encoder lr 2e-5 · CT2 default0.1950.2250.1190.2390.138
turbo frozen encoder lr 1e-5 · CT2 default0.2120.2440.1280.2660.148
v0: frozen encoder, 1,000 steps, unfiltered data0.2540.2840.1470.3050.168
baseline ivrit-ai yi-whisper-large-v3 (no fine-tuning)0.4960.5310.2740.5340.275
baseline ivrit-ai yi-whisper-large-v3-turbo (no fine-tuning)0.5420.5710.2860.5690.283

Accuracy on the product years: the delivered transcripts against the hanachos

With the PDFs in hand, the 1,005 machine transcripts of 5778–5781 (third pass, production model) can be scored against the printed text for the first time. These are whole-file numbers over all days, including days whose printed text is a different sicha or a heavily edited one, so they are upper bounds; the fixed test set (0.081) sits in the same range.

yeardaysreference wordsWERCERmedian day WERreference text
5778247319,4040.0720.0410.059decoded PDF (David + bold crib)
5779266341,9730.0700.0400.057decoded PDF
5780245306,5330.0730.0420.057decoded PDF
5781247312,7400.0900.0500.071clean text layer; 0.080 without the 2 mismatched days

A decoder failure the earlier audit missed. Comparing word counts exposed 14 transcripts far shorter than their text: the silero voice-activity detector had kept 3 windows of a 638 s loud, clipped recording (31 words for 011 18-Tishrei 5781) and dropped 11–52% of the audio in 13 more files, while the "0 empty transcripts" check passed. Lower VAD thresholds recover little; decoding every gap as loud as the speech recovers it but hallucinates on gaps that are crowd noise or singing. The decoder now (1) falls back to fixed windows when the VAD covers under 30% of a recording and (2) by default adds loud gaps as extra windows and keeps each only when its decode is confident (mean token log-prob above −0.35 and at least 1 word/s). Validation: identical output on normal files (control set 0.0955 = 0.0955; the 20-day test set 0.0814 vs 0.0814), 0.231 → 0.135 on the 13 affected files. The ten truncated transcripts were replaced; the originals are kept in output_superseded/.

Where the errors are now

v1 turbo lr 1e-5, CT2 default suppression (test 0.110)v2 turbo lr 3e-5, GPU path (test 0.087)
0 0.2 0.4 0.6 error rate of reference words seen 100+ times · v1 turbo lr 1e-5, CT2 default suppression: 0.055 0.055 seen 100+ times · v2 turbo lr 3e-5, GPU path: 0.044 0.044 seen 100+ times 82.5% of words seen 10–99 · v1 turbo lr 1e-5, CT2 default suppression: 0.153 0.153 seen 10–99 · v2 turbo lr 3e-5, GPU path: 0.114 0.114 seen 10–99 12.5% seen 1–9 · v1 turbo lr 1e-5, CT2 default suppression: 0.343 0.343 seen 1–9 · v2 turbo lr 3e-5, GPU path: 0.274 0.274 seen 1–9 3.8% never seen (OOV) · v1 turbo lr 1e-5, CT2 default suppression: 0.512 0.512 never seen (OOV) · v2 turbo lr 3e-5, GPU path: 0.502 0.502 never seen (OOV) 1.2% abbreviations (gershayim) · v1 turbo lr 1e-5, CT2 default suppression: 0.303 0.303 abbreviations (gershayim) · v2 turbo lr 3e-5, GPU path: 0.175 0.175 abbreviations (gershayim) 6.0% (overlaps)
Error rate of reference words by how often the word occurs in the clean training text (substitutions + deletions on 42,368 dev+test words), plus abbreviations written with gershayim. Frequent words are right 95% of the time; words never seen in training are wrong half the time. The abbreviation class fell from 30% to 17.5% once the decoder could emit the quote and the model improved; what remains there is mostly abbreviation-vs-spelled-out variation that the references themselves are inconsistent about (עאכו"כ vs על אחת כמה וכמה, בנוגע vs בהנוגע). The single largest confusion is דער/די/דעם (gender and case of the article).

v5: the enlarged dataset and the runs on it (in progress)

Status at the snapshot: the fleet has aligned 5781 and is at 1,170 of the remaining 1,511 days (5775–5780, wave 2; ~6 days per minute). When it finishes, every node and this Mac rebuild the dataset from the same day records, the production model scans the train split, the filter drops clips with CER > 0.6, and a guarded launcher starts the runs below unless the cleaned manifest is small or a year lost more than 25% of its clips. Rules kept: data/splits.json unchanged (the 20 dev and 20 test days stay in 5786–5787); the new years go to train only, except test_b = 20 evenly spaced 5781 days (clean text layer) that are never trained on and are reported next to the main test set from now on. Expected size: about 66,000 clips, ~370 h (30,212 clips / 168 h in v4).

runnoderecipewhyexpected
turbo-full-v5-cont-lr1e5-2epbigmac01production model continued 2 ep on all data, lr 1e-5cheapest read of what the new years add~12 h
turbo-full-v5-lr3e5-2epbigmac02production recipe from the base on all data, 2 ep, lr 3e-5same optimizer steps as the 4-epoch v4 run on 2.5× the clips~12 h
large-full-v5-lr2e5-2epsupermac01 (512 GB)large-v3 from the base on all data, 2 ep, lr 2e-5best clip model in v4 (0.083); the target model~35 h on a Mac Studio; hours on the GB10 with CUDA/bf16

v4 verdict, complete (all six runs scored): on the old data nothing beats the production recipe; the seed repeat (0.080 vs 0.081 whole-file) puts run-to-run noise at ~0.001, so table differences under 0.003 mean nothing. Higher learning rates (4e-5, 5e-5) are worse; self-training never wins; large-v3 at lr 3e-5 (0.086 clips / 0.081 whole-file) equals the turbo on whole files at 2.5× the cost and does not reach large-v3 at lr 2e-5 on clips. The v5 data is the only remaining lever.

Next steps

  1. Finish the v5 data (today). Alignment wave 2 → dataset → production-model scan → filter. Read the per-year drop rate: a year above ~10% means a decode or alignment problem in that year's text, not bad audio.
  2. Train the three v5 runs (Thu–Sat) and score each on the fixed test set through the GPU path (whole-file WER is the ranking metric) and on test_b. Decide by whole-file test WER; anything under 0.003 apart is a tie.
  3. If a v5 model wins: a fourth Phase 7 pass with it and the auto windows over 5775–5781 (the project owner's go-ahead needed; the third pass stands until then), then the per-year drafts-vs-hanacha table again.
  4. Attack the remaining errors where they are: half of the remaining word errors are words never seen in training (names, sources, loshon-kodesh); the text-only years (5768–5770, 5774, 5782: ~1.4 M words) are material for a spelling/normalization layer or decoder biasing, and for consistent abbreviation forms.
  5. Compute. The large-v3 runs take 30+ h per Mac Studio; the GB10 (CUDA, bf16) would cut that to hours once its GPU memory and disk are freed. Fleet alignment should move from static shards to a shared queue (the tail idled the fleet ~1 h per wave).
  6. Product and review. Synced subtitles (an SRT per file exists) are the preferred first deliverable; nothing is published or made searchable until Rabonim/Mashpi'im and the hanachos' author have reviewed; ask The Daily Sicha team for the 5782 recordings (text exists) and confirm the two 5781 days whose printed text does not match the audio (17 Kislev, 9 Teves).
  7. Housekeeping. Report every number from the GPU path only (faster-whisper long-form loses ~3 points and suppresses the gershayim by default); keep output_superseded/ and the windowing reports as the audit trail.

Reproducing any number on this page

artifactpath (under Sichos/asr/)notes
clip manifestsdata/manifests/all.jsonl, train.jsonl, train.clean.jsonl, dev.jsonl, test.jsonlone JSON row per clip: audio path, start/end, duration, text, text_raw, day, split, source_year, timing_source, anchor_rate
pseudo-labelsdata/manifests/train.pseudo.jsonl, train.clean_pseudo.jsonlbuilt by build_pseudo_dataset.py + merge_pseudo_manifest.py
splitsdata/splits.jsonrule text, day lists, per-day detail (publication year, month, source year)
day recordsdata/days/<year>/*.jsonAPI payload, offsets, site timings, own alignment (sync_local) with diagnostics
clipsdata/segments/<year>/<day>/NNN_<start×100>.flac16 kHz mono FLAC, on every fleet node
evaluation reportsreports/v1/, reports/aq/, reports/hf/, reports/sweep/per-clip ref/hyp pairs; compare_runs.py recomputes every table from them
training logs and configs~/Sichos/asr/reports/train-<run>.log, runs/<run>/train_config.json on the nodeseval_wer every 500 steps
hanachos from PDFsdata/pdf_text/<year>/NNN.json, _summary.json, _fontmap.jsondecoded day texts, per-font maps and in-vocabulary shares; reports/fleet/pdf-*.log
drafts vs hanachosreports/drafts-vs-pdf/by-year.md, reports/windowing/score_drafts_vs_pdf.py; the windowing comparison (summary.txt, replaced.json)
error analysisreports/error-analysis-*.mdpipeline/error_analysis.py
narrative and decisionsAGENT_HANDOFF.md, README.md, cluster/README.mddated log of every choice under CURRENT POSITION