Neural machine translation when the data barely exists
Dongxiang has almost no parallel data, so we fine-tuned Meta’s NLLB-200 instead of prompting a general model. Cleaning the bilingual corpus, checking subword fertility, registering a new language tag, and why Adafactor beat AdamW on one GPU.