Qwen3.8 27B
Three DeltaNet layers for every attention layer. The model everything was tuned on.
nuclis is an inference engine for Apple Silicon, written from scratch in Zig and Metal. Hand it a 27-billion-parameter model and a failing test suite, and watch it find the bug on your laptop, offline. Every layer was proven against a reference before a single kernel was tuned.
For macOS on Apple Silicon. Free and open source.
A token's trip through Qwen3.8 crosses 64 layers. Before any of them was made fast, each one ran beside a pinned llama.cpp on the same input, on the CPU and on the GPU, until their answers agreed. The reference is a witness, never a source: nothing was copied from it.
16 attention layers remember every token they have seen. Watch their cache grow.
48 Gated DeltaNet layers fold the past into a state that never grows, at token ten or token thirty thousand.
A file it cannot run exactly is turned away at the door. Nothing is quietly requantized.
An M4 Pro with 48 GB, and the same tokens fed to nuclis and to llama.cpp. Qwen3.8 is where the tuning went, and it shows. The other families still run on kernels written for Qwen; their gap is the to-do list.
128 greedy tokens, the warm mean of three runs, release builds on both sides. How these were measured
A small drafter guesses the next few tokens. The big model checks them all in a single pass and keeps exactly what it would have written anyway. Same words, sooner.
An illustration, not a measurement. The drafter guesses the bug; the model overrules it.
On the agent's own task list, Qwen3.8 spends 38% less time in the model. Gemma 4 26B-A4B sits this one out: with 128 experts, checking a guess costs more than it saves.
Each model in the catalogue is pinned to a repository, a commit and a SHA-256. You get the exact bytes that were measured, or you get an error. Every one reads images and brings its own drafter.
Three DeltaNet layers for every attention layer. The model everything was tuned on.
Trained knowing it would be quantized, so every matrix ships at 4 bits.
A mixture of experts: 128 of them, 8 awake for any one token.
Built for small devices, with per-layer embeddings and layers that share a cache.
Dense, paired with a block drafter that guesses a run of tokens at once.
If its architecture has an adapter, it runs: finetunes, other quantizations, even a ternary Bonsai. nuclis model inspect tells you before you download 16 GB.
nuclis agent works in the folder you start it in, with six tools and no more. The host sets every limit: steps, output, results. When something is cut short it says so, and when a command fails the model reads the failure and tries again.
bashruns a command, on a leashreadpages through a filegrepfinds it, grouped by fileglobfinds files by nameeditchanges lines, shows the diffwritestarts a new filenuclis serve speaks OpenAI's Chat Completions, so an agent or a program written for an OpenAI-compatible server changes only its base URL. Tools, images and the reasoning effort map onto the model's own template, replies stream, and the server keeps what a conversation has already sent: an outside coding agent's fifth request reused 10,419 of its 10,591 prompt tokens.
chat/completionstools, images, streamedembeddingsvectors, batched on the GPUdecisionstyped answers, one state or manysystemonea Jev client, unchangedmodelswhat is open, what could behealthis it up$ nuclis serve --chat-model qwen3.8-27b
# then point any OpenAI-compatible client at
http://127.0.0.1:9000/v1
Not every question needs a paragraph back. nuclis decide asks a model yes or no, which one, or how much, and gets the answer in one pass without writing a word. Laya reads a 500-token log in a fifth of a second. Cloudflare's clef-flash takes on long states and pictures. nuclis serve keeps both warm behind the API.
$ nuclis decide --noul 'Does this log need a person now?' \
--state-file disk.log --state-file payments.log
nuclis embed runs Google's EmbeddingGemma 2: a sentence, a picture, a spoken clip, or all three in one input become 768 numbers. Things that mean the same land close together, whatever they are made of, so one cosine compares anything with anything. On the BF16 file the vectors match Google's own pipeline to 1 − cos ≤ 5e-11. nuclis makes them; your search keeps the index.
text64 passages of 256 tokens at 38 a secondimageabout half a second a pictureaudio4.8 s of speech in 0.13 s$ nuclis embed --image photo.jpg --audio memo.m4a "the shopping list"
$ curl -s localhost:9000/v1/embeddings -d '{"input": ["a cat", "a kitten"]}'
Unpack it, let macOS trust it, pull a model about the size of a long film. The binary is not notarized, so macOS holds it until the third line says otherwise. Every release is attested: gh attestation verify shows where it was built.
You need Zig 0.17.0 and the Command Line Tools. make metal builds the release binary with the Metal backend.
# unpack the release
tar xzf nuclis-v0.5.0-aarch64-macos.tar.gz
cd nuclis-v0.5.0-aarch64-macos
xattr -d com.apple.quarantine nuclis
# fetch a model, then put it to work
./nuclis model pull qwen3.8-27b --all
./nuclis agent
git clone https://github.com/tildaslashalef/nuclis
cd nuclis && make metal
./zig-out/bin/nuclis model pull qwen3.8-27b --all
./zig-out/bin/nuclis config init --discover
./zig-out/bin/nuclis agent
nuclis began as a way to learn how inference really works, from the bytes on disk to the sampled token, and to learn Zig along the way. AI coding agents write much of the code. The rules they follow are public, every piece of work is logged with its evidence, and every number on this page was measured, never estimated.