Routing a request in six languages, and knowing when to hand it to a person
Why
A support assistant needs to route a request to the right intent in the customer’s own language, and know when the request is something it was never trained on. I wanted a real number for both, in English and five Indian languages, using a method small enough to fine-tune on a laptop: how much do the other five languages cost when the model only ever saw English, how much of that cost does training on all six recover, and can the model’s own confidence catch the requests it should hand to a person.
How
I fine-tuned a small multilingual encoder with LoRA, training a fraction of one percent of its parameters, on a public benchmark of voice-assistant commands in English, Hindi, Tamil, Telugu, Kannada and Malayalam. Six intents were held out before any training so the study could test hand-off on real unseen requests, drawn by a seed fixed in advance. Every number came from a single sealed test pass opened once, against predictions written down before any model existed.
What came out
Training on English only costs 15 to 28 points of accuracy on the other five languages; training on all six recovers 83% of that on average, closing Hindi’s 15-point gap to 2.7 and Tamil’s 27.7-point gap to 6.8. Hand-off to a person works in English (0.792 AUROC) but is weaker in the other languages when the model only trained on English (0.611 to 0.702), and a threshold tuned to keep 95% of legitimate English requests keeps only 76% to 85% of legitimate requests in the Indian languages until the model also trains on them.
What broke
The comparison the whole study was framed around, whether fine-tuning beats a frozen model, could not be answered: the frozen baseline failed its own sanity check against a published reference number, so under the rule fixed in advance that entire arm is void, reported but not used. A distance-based hand-off score expected to beat a plain confidence score by three points did not, landing two thousandths of a point ahead of it, statistically nothing. And a plain character n-gram model with no neural network at all matched the fine-tuned model within about a point in every language once it also saw that language in training, so most of the multilingual routing problem was never really about the fine-tuning.
Mind the gap
Private. The results above are final under the pre-registration; there is no public repository to link.