nuclis

Right first. Then fast.

nuclis is an inference engine for Apple Silicon, written from scratch in Zig and Metal. Hand it a 27-billion-parameter model and a failing test suite, and watch it find the bug on your laptop, offline. Every layer was proven against a reference before a single kernel was tuned.

For macOS on Apple Silicon. Free and open source.

A real session, retyped: Qwen3.8-27B on an M4 Pro with speculative decoding on. The long pauses are trimmed; the rest runs at the speed it was recorded.

Every layer, cross-examined.

A token's trip through Qwen3.8 crosses 64 layers. Before any of them was made fast, each one ran beside a pinned llama.cpp on the same input, on the CPU and on the GPU, until their answers agreed. The reference is a witness, never a source: nothing was copied from it.

 

 

16 attention layers remember every token they have seen. Watch their cache grow.

48 Gated DeltaNet layers fold the past into a state that never grows, at token ten or token thirty thousand.

Bit for bit
Every quantized block decodes to exactly what the reference decodes. Not close: exactly.
Layer for layer
The real model on real prompts. Every layer's output and the final logits compared on both backends, inside a tolerance written down in advance.
Page for page
Four thousand tokens of a fixed text, scored within half a percent of the reference's perplexity.

A file it cannot run exactly is turned away at the door. Nothing is quietly requantized.

Honest numbers, one machine.

An M4 Pro with 48 GB, and the same tokens fed to nuclis and to llama.cpp. Qwen3.8 is where the tuning went, and it shows. The other families still run on kernels written for Qwen; their gap is the to-do list.

Tokens per second

nuclis llama.cpp

128 greedy tokens, the warm mean of three runs, release builds on both sides. How these were measured

Guess ahead, check in bulk.

A small drafter guesses the next few tokens. The big model checks them all in a single pass and keeps exactly what it would have written anyway. Same words, sooner.

An illustration, not a measurement. The drafter guesses the bug; the model overrules it.

Decode, drafter off and on

off on

On the agent's own task list, Qwen3.8 spends 38% less time in the model. Gemma 4 26B-A4B sits this one out: with 128 experts, checking a guess costs more than it saves.

Pinned to the byte.

Each model in the catalogue is pinned to a repository, a commit and a SHA-256. You get the exact bytes that were measured, or you get an error. Every one reads images and brings its own drafter.

Qwen3.8 27B

Three DeltaNet layers for every attention layer. The model everything was tuned on.

qwen3.8-27b16.5 GB

Gemma 4 12B QAT

Trained knowing it would be quantized, so every matrix ships at 4 bits.

gemma-4-12b-qat6.7 GB

Gemma 4 26B-A4B

A mixture of experts: 128 of them, 8 awake for any one token.

gemma-4-26b-a4b14.2 GB

Gemma 4 E4B QAT

Built for small devices, with per-layer embeddings and layers that share a cache.

gemma-4-e4b-qat4.2 GB

Muse Glimmer 30B

Dense, paired with a block drafter that guesses a run of tokens at once.

muse-glimmer-30b15.9 GB

Something else?

If its architecture has an adapter, it runs: finetunes, other quantizations, even a ternary Bonsai. nuclis model inspect tells you before you download 16 GB.

An agent that knows its limits.

nuclis agent works in the folder you start it in, with six tools and no more. The host sets every limit: steps, output, results. When something is cut short it says so, and when a command fails the model reads the failure and tries again.

  • bashruns a command, on a leash
  • readpages through a file
  • grepfinds it, grouped by file
  • globfinds files by name
  • editchanges lines, shows the diff
  • writestarts a new file

Bring your own agent.

nuclis serve speaks OpenAI's Chat Completions, so an agent or a program written for an OpenAI-compatible server changes only its base URL. Tools, images and the reasoning effort map onto the model's own template, replies stream, and the server keeps what a conversation has already sent: an outside coding agent's fifth request reused 10,419 of its 10,591 prompt tokens.

  • chat/completionstools, images, streamed
  • embeddingsvectors, batched on the GPU
  • decisionstyped answers, one state or many
  • systemonea Jev client, unchanged
  • modelswhat is open, what could be
  • healthis it up
$ nuclis serve --chat-model qwen3.8-27b
# then point any OpenAI-compatible client at
http://127.0.0.1:9000/v1

Questions with typed answers.

Not every question needs a paragraph back. nuclis decide asks a model yes or no, which one, or how much, and gets the answer in one pass without writing a word. Laya reads a 500-token log in a fifth of a second. Cloudflare's clef-flash takes on long states and pictures. nuclis serve keeps both warm behind the API.

$ nuclis decide --noul 'Does this log need a person now?' \
    --state-file disk.log --state-file payments.log
Decision Dungeons A playground of seeded worlds where a decision model makes every call, and every call is scored. Play it

Anything, as a vector.

nuclis embed runs Google's EmbeddingGemma 2: a sentence, a picture, a spoken clip, or all three in one input become 768 numbers. Things that mean the same land close together, whatever they are made of, so one cosine compares anything with anything. On the BF16 file the vectors match Google's own pipeline to 1 − cos ≤ 5e-11. nuclis makes them; your search keeps the index.

  • text64 passages of 256 tokens at 38 a second
  • imageabout half a second a picture
  • audio4.8 s of speech in 0.13 s
$ nuclis embed --image photo.jpg --audio memo.m4a "the shopping list"
$ curl -s localhost:9000/v1/embeddings -d '{"input": ["a cat", "a kitten"]}'

From download to first token.

Unpack it, let macOS trust it, pull a model about the size of a long film. The binary is not notarized, so macOS holds it until the third line says otherwise. Every release is attested: gh attestation verify shows where it was built.

# unpack the release
tar xzf nuclis-v0.5.0-aarch64-macos.tar.gz
cd nuclis-v0.5.0-aarch64-macos
xattr -d com.apple.quarantine nuclis
# fetch a model, then put it to work
./nuclis model pull qwen3.8-27b --all
./nuclis agent

Built in the open, receipts included.

nuclis began as a way to learn how inference really works, from the bytes on disk to the sampled token, and to learn Zig along the way. AI coding agents write much of the code. The rules they follow are public, every piece of work is logged with its evidence, and every number on this page was measured, never estimated.