AIToday
Large Language ModelsOpen-Source AIHacker NewsPublished: May 5, 2026, 10:00 JST1 min read

SMILE Serve – a production-ready inference server on Quarkus supporting Classic ML, ONNX, and Llama 3 chat models with bundled React web UI

SMILE Serve – a production-ready inference server on Quarkus supporting Classic ML, ONNX, and Llama 3 chat models with bundled React web UI

3 Key Points

  1. SMILE Serve bundles three inference capabilities on the JVM: Classic ML (serialized SMILE models), ONNX Runtime (open-format models), and LLM Chat (Llama 3 chat completions with conversation persistence), each accessed via separate API prefixes.

  2. Deployment options include Docker (quickest start), JAR packaging, Uber-JAR (single self-contained file), and native executable compilation via GraalVM for sub-millisecond startup and lower memory use.

  3. Classic ML and ONNX models are auto-discovered from mounted directories at startup; the Chat API supports a configurable maximum sequence length of 4096 tokens and can be backed by PostgreSQL for chat history persistence.

Ask the AI about this article →

Get the latest Large Language Models news every morning

For example, today's edition would include:

  • Visko raises $10M, launches live AI video model OrbisSiliconANGLE AI · 2h ago
  • Runway unveils Solaris, an AI that generates app interfaces in real timeTHE DECODER · 2h ago
  • Google AI Search flags Facebook users as dangerTHE DECODER · 2h ago

AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.

Free · takes 30 seconds · unsubscribe anytimeWhat is AIToday? →

Ask AI

Ask AI anything about this article. Q&As are published on this page for other readers too.

Related Articles

Next articleGreg Brockman Defends $30 Billion OpenAI Stake During Musk Trial, Says He Poured 'Blood, Sweat, and Tears' Into Company