AIToday
Large Language ModelsAI in HealthcareAI Safety & AlignmentarXiv cs.CLPublished: Apr 17, 2026, 13:00 JST1 min read

Researchers achieve 32.87 score on clinical QA task using two-stage QLoRA fine-tuning of Qwen3-4B model

Researchers achieve 32.87 score on clinical QA task using two-stage QLoRA fine-tuning of Qwen3-4B model

3 Key Points

  1. QU-NLP team applies two-stage Quantised Low-Rank Adaptation (QLoRA) to Qwen3-4B loaded in 4-bit NF4 quantisation for the ArchEHR-QA 2026 shared task

  2. System first trained on 30,000 samples from emrQA-MedSQuAD corpus for clinical domain knowledge, then on 20 annotated development cases for task-specific output

  3. Subtask 3 (answer generation) achieves overall score of 32.87 with BLEU=9.42, ROUGE-L=27.04, SARI=55.42, BERTScore=43.00, and MEDCON=37.04

  4. Subtask 4 (evidence alignment) uses weighted ensemble of BM25, TF-IDF, and fine-tuned cross-encoder to reach 67.16 micro-F1 score on 100-case test set

Ask the AI about this article →

Get the latest Large Language Models news every morning

For example, today's edition would include:

  • DataAgent launches with $10M to auto-fix Kubernetes faultsSiliconANGLE AI · 45m ago
  • SK Hynix custom HBM boosts inference up to 5.15xDIGITIMES Asia · 45m ago
  • Nvidia Earnings: Boring by Design, Avoiding a Consolidated WorldStratechery (Ben Thompson) · 45m ago

AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.

Free · takes 30 seconds · unsubscribe anytimeWhat is AIToday? →

Ask AI

Ask AI anything about this article. Q&As are published on this page for other readers too.

Related Articles

Next articleNew AI agent called AIBuildAI automates the entire machine learning development process, from architecture design to implementation and debugging.