Loading MicroLLM lab…

Initializing WebGPU engine & model catalog…

On-Device AI Lab WebGPU · Zero-Server · Q4 Quantized

MicroLLM lab

Run, benchmark, and compare Small Language Models (SLMs) directly in your browser with hardware-accelerated WebGPU — 100% private, zero server cost, zero accounts.

🎯 Why Small Language Models?

Frontier models (GPT-4, Claude) cost millions to serve and add network latency. Compact models (25M–360M) operate as an ultra-fast, efficient edge layer:

  • Edge Computing & Privacy: Runs entirely on your phone or laptop. No prompt or user data ever leaves your device.
  • Fast Triage & Task Routing: Classify queries, filter spam, and extract intent in milliseconds to decide if an expensive cloud LLM is even needed.
  • Zero Cloud Cost: Infinite concurrency powered by client GPUs with no API bills.
  • Ultra-Low Latency: Sub-10ms time-to-first-token for instantaneous autocomplete and real-time agents.

🔍 Technical Vocabulary

Small Language Model (SLM)
A compact neural model (~25M–360M parameters) designed for efficient edge intelligence and specific tasks rather than broad trivia.
Q4 (4-bit Quantization)
Compressing weights from 16-bit floats to 4 bits per parameter, shrinking memory footprint by 75% so 100M+ models fit in ~50–84 MB of browser memory with near-lossless generation quality.
WebGPU
The modern W3C standard API that executes compute shaders directly on your hardware GPU (Apple Silicon Metal, DirectX 12, Vulkan) right inside your browser window.
How to use this lab in 3 steps:
1
Load a model: Click Load on any model card below. It caches directly into your browser's private IndexedDB (no file download to disk).
2
Chat on-device: Select your loaded model and scroll down to the Chat panel to test prompts and watch live tok/s speed.
3
Benchmark & Compare: Switch to the Benchmarks tab to run objective speed/accuracy tests and generate your shareable performance certificate.

Loaded on this device: … (Saved in browser IndexedDB cache)

Next action:

Objective checks (regex / exact tokens), not writing quality. A 135M model is allowed to fail — that is the measurement. Pick models, then run. Estimate uses your last tok/s if we have one.

🏎️ Highest Peak Speed — Fastest single test
⚡ Highest Sustained Speed — Continuous 256-tok decode
📊 Avg Sustained Speed — Across tested models
⏱️ Total Benchmark Score — Cumulative suite wall time
Next action:

Speed (tokens/s, sustained decode, suite wall) and accuracy (pass rate on objective tests) from runs in this browser. Numbers stay on this machine. Charts use the latest suite per model.

🏎️ Highest Peak Speed — Fastest single test
⚡ Highest Sustained Speed — Continuous 256-tok decode
📊 Avg Sustained Speed — Across loaded models
⏱️ Total Benchmark Score — Total execution wall time

Verified Benchmark Certificate & Social Share

Generate and download a verifiable performance certificate with your device hardware, peak and sustained tokens/second, and share your score.

MicroLLM Lab Benchmark Certificate

Next action:

Write a benchmark in JavaScript

The editor is eval()’d in this origin, then each check runs on the model’s decoded text.

Prompt for a larger LLM

Function form


            
Next action: