
SMILE Serve bundles three inference capabilities on the JVM: Classic ML (serialized SMILE models), ONNX Runtime (open-format models), and LLM Chat (Llama 3 chat completions with conversation persistence), each accessed via separate API prefixes.
Deployment options include Docker (quickest start), JAR packaging, Uber-JAR (single self-contained file), and native executable compilation via GraalVM for sub-millisecond startup and lower memory use.
Classic ML and ONNX models are auto-discovered from mounted directories at startup; the Chat API supports a configurable maximum sequence length of 4096 tokens and can be backed by PostgreSQL for chat history persistence.
Ask the AI about this article →
For example, today's edition would include:
AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.
Free · takes 30 seconds · unsubscribe anytimeWhat is AIToday? →
Ask AI anything about this article. Q&As are published on this page for other readers too.
Visko raised $10 million in pre-seed funding from Llama Ventures and opened public access to its first foundat…
AI company Runway has unveiled Solaris, the first model in a new category it calls "Interface World Models." I…

Google's AI search gave advice to call emergency services for users alone with an African, Indian, or Pakistan…

John Deere introduced JD, a conversational AI tool that lets farmers ask open-ended questions about their hist…

Nvidia CEO Jensen Huang said on Fox Business that AI is creating 'hundreds of thousands' of jobs, including in…

Israeli startup DataAgent Ltd