AIToday
Large Language ModelsOpen-Source AIQiita 機械学習Published: Oct 6, 2026, 19:01 JST

Laya test: synthetic data lifts 56%, hurts 30%

Laya test: synthetic data lifts 56%, hurts 30%

3 Key Points

  1. What happened

    Testing the open-source Laya decision model on routing questions to one of ~1,000 database tables, fine-tuning on synthetic queries raised validation accuracy 56% but cut real-conversation accuracy 30%.

  2. Why it matters

    The synthetic queries were short keyword phrases, while real users send long conversational questions, so more training pulled the model away from real traffic—and the table-held-out validation set missed it, since its queries shared the training style.

  3. What to watch

    The author still sees value in Laya as a confidence gate, noting correct answers showed median confidence of 0.88–0.90 versus 0.68–0.73 for wrong ones, so the fix hinges on validation sets that vary query style.

WHO IT HITSTeams building their own local query-routing or table-selection layers on small open models are the ones affected. The result suggests that engineers validating such models on metadata-derived synthetic queries may see accuracy gains that do not carry over to real users, so they should check validation data for style differences from production traffic.

Not sure about something? Ask the AI

Questions and answers are published on this page.

Summaries like this, in your inbox every morning.

Context & Analysis

The experiment started from a practical constraint: routing free-form questions to the right table among roughly 1,000, fast, on local hardware, without shipping schema metadata to an external API. Laya looked like a clean fit on paper, since it is a non-autoregressive model that takes a state and a typed question with a defined set of options and returns a calibrated choice in one forward pass. The 512-token input limit, however, pushed the setup into a two-stage design: retrieve about 15 candidates, then let Laya decide among them.

Retrieval turned out to set the ceiling, and lexical search struggled with schema vocabulary, where users say "daily sales by store" but the table is named XQR7_DTV_AGG. A small embedding model closed much of that gap cheaply, while roughly 30 hand-written abbreviation expansions added only one point. The author also found that concatenating the last two user utterances roughly doubled every metric, because follow-up questions cannot be resolved without history.

Fine-tuning was meant to be the payoff, and the documentation was free: splitting "typical phrasing" strings from each table's docs yielded about 25,000 query-table pairs. Training on a 56-thread Xeon took about 17 hours per epoch, close to three days for four epochs. Validation accuracy rose 56%, but accuracy on the 115 real conversation turns fell 30%, because the model learned the terse keyword style of the synthetic data rather than conversational traffic. The table-held-out validation set never flagged this, since its queries came from the same synthetic distribution. The author's recommended path is to fix the training data's style rather than the model, to invest in retrieval before the decision stage, and to use Laya mainly as a confidence gate that answers locally when sure and escalates to a larger LLM when not.

FAQ
How was Laya tested in this experiment?
The author used it to pick one of roughly 1,000 database tables for free-form user questions, with hybrid retrieval first narrowing the list to 15 candidates. Each candidate's ~140-character description was then fed to Laya for the final choice.
What retrieval method found the right table most often?
A hybrid of BM25 and a 33-million-parameter sentence embedding model with Reciprocal Rank Fusion reached hit@15 of 0.65 with context, versus 0.45 for field-weighted BM25 alone. A hand-written abbreviation alias list added just one point.
Did fine-tuning help or hurt Laya in real use?
It raised accuracy on synthetic validation queries from 0.244 to 0.380 after four epochs. But on 115 real conversation turns it fell from 0.200 to 0.139, a 30% drop.
Qiita 機械学習Read Original Article

AI news that matters for your work, delivered every morning.

Pick your industry and the AI tools you use, and get news related to your work every day.

Free · 30 seconds with Google · unsubscribe anytimeWhat is AIToday? →

Ask AI

Ask AI anything about this article. The AI reads this article, earlier AIToday articles, and Wikipedia, and cites its sources. Q&As are published on this page for other readers too.

Questions and answers are published on this page.

Related Articles

Next articleSheikh Tahnoon: the UAE spymaster betting $1.5tr on AI dominance