
What happened
Testing the open-source Laya decision model on routing questions to one of ~1,000 database tables, fine-tuning on synthetic queries raised validation accuracy 56% but cut real-conversation accuracy 30%.
Why it matters
The synthetic queries were short keyword phrases, while real users send long conversational questions, so more training pulled the model away from real traffic—and the table-held-out validation set missed it, since its queries shared the training style.
What to watch
The author still sees value in Laya as a confidence gate, noting correct answers showed median confidence of 0.88–0.90 versus 0.68–0.73 for wrong ones, so the fix hinges on validation sets that vary query style.
WHO IT HITSTeams building their own local query-routing or table-selection layers on small open models are the ones affected. The result suggests that engineers validating such models on metadata-derived synthetic queries may see accuracy gains that do not carry over to real users, so they should check validation data for style differences from production traffic.
Summaries like this, in your inbox every morning.
The experiment started from a practical constraint: routing free-form questions to the right table among roughly 1,000, fast, on local hardware, without shipping schema metadata to an external API. Laya looked like a clean fit on paper, since it is a non-autoregressive model that takes a state and a typed question with a defined set of options and returns a calibrated choice in one forward pass. The 512-token input limit, however, pushed the setup into a two-stage design: retrieve about 15 candidates, then let Laya decide among them.
Retrieval turned out to set the ceiling, and lexical search struggled with schema vocabulary, where users say "daily sales by store" but the table is named XQR7_DTV_AGG. A small embedding model closed much of that gap cheaply, while roughly 30 hand-written abbreviation expansions added only one point. The author also found that concatenating the last two user utterances roughly doubled every metric, because follow-up questions cannot be resolved without history.
Fine-tuning was meant to be the payoff, and the documentation was free: splitting "typical phrasing" strings from each table's docs yielded about 25,000 query-table pairs. Training on a 56-thread Xeon took about 17 hours per epoch, close to three days for four epochs. Validation accuracy rose 56%, but accuracy on the 115 real conversation turns fell 30%, because the model learned the terse keyword style of the synthetic data rather than conversational traffic. The table-held-out validation set never flagged this, since its queries came from the same synthetic distribution. The author's recommended path is to fix the training data's style rather than the model, to invest in retrieval before the decision stage, and to use Laya mainly as a confidence gate that answers locally when sure and escalates to a larger LLM when not.
Pick your industry and the AI tools you use, and get news related to your work every day.
Free · 30 seconds with Google · unsubscribe anytimeWhat is AIToday? →
Ask AI anything about this article. The AI reads this article, earlier AIToday articles, and Wikipedia, and cites its sources. Q&As are published on this page for other readers too.
Anthropic said that from October 6 onward, Cowork tasks on the Claude Pro and Max plans run in the cloud, and…

Cohere launched North 2, a revamped enterprise platform with orchestration that coordinates AI agents tapping…

Oracle Japan's president Tomomitsu Misawa introduced "Oracle Fusion Claw," an enterprise take on the desktop-a…

A Mat3ra notebook compared graphene on Ni(111) across four registries (fcc, hcp, hollow, bridge) using MACE-MP…

KDDI's ELYZA said on October 2 it set up "ELYZA RSI Research" to study recursive self-improvement, and CEO Yuy…

On a recent episode of The AI Investor Podcast, Eric Bleeker and Austin Smith broke down Google's release of G…
