A small language model (SLM) is a model small enough to run on a phone, a school laptop, or a single GPU next to your API — typically under about 8 billion parameters after quantisation. We use them when the feature must work offline, stay inside a country, or stay cheap at millions of calls a day.
Cloud frontier models are still better at open-ended reasoning. They are also slower, more expensive, and they leave the building. For EduKidGames hints, Polylingo pronunciation prompts, and on-device grammar checks, that trade is the wrong way around. The child in a classroom with a weak connection should not wait on a region outage in Virginia.
What we keep in the cloud
Curriculum authoring, long-form lesson planning, and anything that needs fresh web knowledge stays on a hosted model with retrieval. The SLM on the device handles the tight loop: classify the learner’s answer, suggest the next card, or rewrite a sentence at a given CEFR level. Those tasks have a narrow output shape. Narrow tasks are where small models stop looking “worse” and start looking faster.
Engineering the box it lives in
On mobile we ship a quantised GGUF or Core ML / ONNX package behind a feature flag, with a size budget called out in the PR the same way we call out an image asset. On the server, an SLM sits beside the ASP.NET process on a dedicated inference host — not inside the web node — and we treat it like any other dependency: health check, queue, timeout, fallback.
Fallback matters. If the on-device model cannot classify a free-text answer with enough confidence, we degrade to a rule or a short cloud call, and we log the miss. That log is how we decide whether to fine-tune or just rewrite the prompt.
Fine-tune last
Teams jump to LoRA because it feels like “real AI”. Most of our wins came from better labels and a smaller task. Fine-tune when the error is systematic and you have a few thousand gold examples from real learners — not a weekend of synthetic chat.
Questions we keep getting
Will SLMs replace GPT-class APIs? Not for open work. They replace the 80% of calls that are classification, rewriting, and routing.
Do parents notice? They notice battery, latency, and whether the app works on the school bus. They do not notice parameter counts.
What stack do you run? ONNX Runtime or llama.cpp on the edge; a small GPU box behind our .NET APIs when we need a studio-hosted model for EU data.