Back
Kevin Riedl

20 min read · 14 Jul 2026
Last reviewed

Next
Made on your device, with no Instagram connection. We copy the post link for Instagram’s Link sticker.

Colibri Runs GLM-5.2 on Consumer Hardware. Here Is the Catch.

Colibri can still run GLM-5.2 on roughly 25 GB of RAM. That remains a capacity breakthrough, not a promise of fast chat. The original developer measurement is 0.05-0.1 generated token per second from a cold cache. At those historical rates, 100 output tokens take about 17-33 minutes, excluding prompt processing. Later container-specific runs are faster, so this is not a universal ceiling for a 25 GB budget.

Updated and source-reviewed on . The latest release verified for this review is Colibri v1.12.0, published on 20 September. The old description of an eight-family, mostly text-oriented experiment is no longer sufficient. Brio can score allowed answers without generating a reply, the dashboard has been redesigned, and several backend and validation fixes have landed. Version 1.11.0 had already added DeepSeek V4.1 Flash as the ninth engine family.

This article keeps its original focus: what Colibri can actually do on your hardware, what it costs in latency and memory, and how to test it without mistaking a successful launch for a production deployment. All speeds below are attributed upstream results, not Wavect measurements.

What is Colibri, and what changed since our previous review?

Colibri, styled Colibrì by its maintainers, is an Apache-2.0 inference runtime that places model weights across NVMe storage, system RAM and optional VRAM. It does not squeeze a 744-billion-parameter model into 25 GB. The checkpoint stays on disk; only the parts needed now must be in faster memory. The tagged README distinguishes the dependency-free C inference engine from the Python launcher and HTTP gateway. “No Python at runtime” is therefore misleading for the normal coli workflow.

Material changes from the article's v1.10.1 baseline
AreaVerified changePractical consequence
Model supportNine engine families; the GLM family also loads GLM-5.3Do not confuse the full GLM-5.3 checkpoint with GLM-5.3-Flash.
Decision scoringBrio through POST /v1/brio, the terminal and dashboardEvaluate closed-set tasks without asking for a generated JSON answer.
GPU pathsQwen3.8 CUDA tier; GLM-5.3-Flash Metal supportOld blanket “CPU-only” statements are stale; verify each engine/backend combination.
Smaller-model performanceQwen3.6 kernel and resident-memory improvementsA speed measured on Qwen3.6 is not a new GLM-5.2 speed.
ReliabilityInput-validation, context-limit, tokenizer, memory and platform fixesRe-run your exact workload after an upgrade rather than only replacing the binary.

The Colibri GLM-5.3 container is approximately 419 GB and uses the same engine family as GLM-5.2. Unlike the recommended 5.2 container, it ships without an MTP head, so that speculative decoding path stays off. This is a different checkpoint choice, not evidence that an existing 5.2 installation automatically becomes 5.3.

How does a 744B model run with 25 GB of RAM?

GLM-5.2 is a Mixture-of-Experts model: about 40B of its roughly 744B parameters are active for a token. Its official model card describes the model, but its model-level capability claims do not establish the quality of a particular Colibri quantization. Colibri keeps roughly 9.9 GB of dense weights resident and fetches selected routed experts from storage. The original layout has 19,456 expert blocks across 75 MoE layers and the MTP head.

A per-layer cache, learned hot-expert pins, operating-system caching and optional GPU placement avoid some repeated reads. The original cold estimate is roughly 11 GB of expert reads per generated token. That is a workload/layout estimate, not a constant for every supported model, every checkpoint or every speculative mode.

There are two separate ideas here. Quantization changes the weight representation; tiering changes where those weights wait. Preserving the same quantized result across CPU and GPU does not prove that quantization preserved the original model's accuracy. Conversely, an exact engine can still be too slow for your application.

Which models does Colibri support, and how much hardware do they need?

Use the v1.12.0 family registry to distinguish an engine family from a checkpoint name. Nine families do not mean nine arbitrary model IDs. Each has architecture-specific routing, attention, tensor formats and chat handling. Colibri is not a generic GGUF loader.

Checkpoint planning guide, not a throughput or fit guarantee
Family / checkpointApproximate disk footprintMemory caveat
GLM-5.2 / GLM-5.3429 GB current GLM-5.2 download / 419 GB GLM-5.3, grouped int4Upstream lists a 16 GB minimum; the original 25 GB result is slow. Cache and context need additional headroom.
GLM-5.3-Flash195 GB convertedAbout 25 GB in the documented configuration; not the full GLM-5.3.
Inkling469 GBThe roughly 25 GB route requires the compressed dense-weight container.
Kimi K31.6 TB32 GB or more in the upstream planning table; verify the actual working set.
DeepSeek V4 Flash167 GB for the full checkpointUpstream lists 16 GB minimum, 32 GB more comfortable. Pruned variants are different checkpoints.
DeepSeek V4.1 Flash510 GBDo not infer a 25 GB fit from other engines; inspect its dense set, caches and planner.
Qwen3.8-Flash-Next185.5 GBPlan from the newer working-set measurements, not the old 16 GB summary.
Qwen3.6-35B-A3BAbout 20 GB, grouped int4A smaller, resident-model candidate; RAM use depends on the current kernel and configuration.
OLMoEAbout 7 GB, int8The upstream entry lists 8 GB RAM; a useful smaller-system test, not GLM-level capability.

Important disk-size correction: the tagged README still calls the GLM-5.2 container 372 GB. The container card has differing size summaries, while its current file listing shows about 429 GB. Budget for the exact revision and selected files, not the old headline; decimal GB and binary GiB should not be mixed. A smaller E8/IQ3 container lists about 289 GB, but changes the representation and has its own accuracy and decoding-cost trade-offs. It is not a free size reduction.

DeepSeek V4.1 Flash's engine documentation explains why total checkpoint size is a poor proxy for per-token work: about 203 GB is disk-backed n-gram memory, and routed experts cost roughly 4.5 GB per token before cache hits. The weights need no conversion, but the documented setup still creates a small dsv41_engram.json sidecar. “No conversion” does not mean “no preparation.”

There is also an important documentation conflict. The README's old Qwen3.8 summary still says 16 GB minimum and CPU-only. The detailed Qwen3.8 measurements put a cap-32 short request at about 16.5 GiB peak RSS and describe 24 GB as a practical host size for that setup, with 32 GB for a larger cache. The v1.12.0 release adds CUDA despite stale “no backend” prose. Prefer the release-specific implementation and measured configuration over an isolated requirements row.

How fast is Colibri on consumer hardware?

The upstream benchmark collection preserves the following GLM-5.2 results. The RAM-capped Core Ultra 9 row comes from the container card linked above. These are historical runs with different builds, cache histories and containers, not a controlled v1.12.0 comparison. Some older rows do not meet the project's newer GPU-correctness reporting requirement.

Reported GLM-5.2 decode speeds; 100-token times are our arithmetic
Reported setupDecode rate100 tokens, decode only
Original WSL2 developer machine, 25 GB RAM0.05-0.1 tok/s cold16.7-33.3 minutes
Container author's Core Ultra 9 / RTX 5080 host, RAM budget capped to 25 GB0.31-0.38 tok/s4.4-5.4 minutes
Mac Mini M4 Pro, 48 GB, Metal0.30 tok/s5.6 minutes
M5 Max, 128 GB, Metal and 46.9 GB learned pin2.06 tok/s48.5 seconds
251 GiB host, six RTX 5090s, full expert residency5.8-6.8 tok/s14.7-17.2 seconds

The six-GPU result is not a consumer-laptop baseline. Later selective NUMA tests reached about 9 tok/s on a specific multi-socket host. Likewise, the release's Qwen3.6 improvement from 12.8 to 15.7 tok/s, with the resident set falling from 29 to 17 GB, belongs to that Qwen workload. It does not make GLM-5.2 a 15.7 tok/s model on 25 GB.

For a real service, measure the entire request: queue wait, prompt prefill, generation and any tool round trips. A reused prompt can have near-zero prefill while a new document is slow. A dashboard screenshot of a warm turn does not establish fresh-document latency.

What should you optimize first: RAM, NVMe or GPU?

Start with the stage consuming wall-clock time. On a small-memory GLM system, repeated expert misses can dominate. On a large resident system, CPU matrix multiplication or memory bandwidth can dominate instead. The historical same-machine SSD comparison raised measured bandwidth from 1.51 to 8.81 GB/s, but generation only from 0.10 to 0.28 tok/s because compute became more important.

Use cold, uncached shard reads rather than the SSD's advertised sequential maximum. On macOS, the documented F_NOCACHE test does not evict pages already cached by a previous run. Also compare multiple prompts: learned pins can favor the workload that created them. A larger cache or more RAM assigned to the engine is not automatically faster if it increases contention or removes system headroom.

A GPU helps when the chosen backend removes the measured bottleneck. It does not remove disk misses by existing in the machine. Keep an output-quality check beside every CUDA, HIP, Metal or Vulkan timing: upstream records a HIP case with similar throughput but dramatically worse perplexity. Speed alone would have missed the defect.

MTP is another experiment, not a universal switch for free speed. GLM-5.2 needs the correct int8 MTP head, and extra speculative expert reads can lose on a cold cache. The router's --topp 0.7 shortcut is an explicitly lossy routing change, not the same as ordinary output-token sampling. Re-evaluate quality whenever you change model semantics.

How much quality does the GLM-5.2 int4 path retain?

The available evidence remains narrower than a broad “frontier quality on a laptop” claim. Upstream's small GLM-5.2 run reports 62.5% mean normalized accuracy across HellaSwag, ARC and MMLU, with only 40 questions per task. Its separate OLMoE experiment measured an 8.2-percentage-point quantization loss, of which grouped scales recovered about 63%. That is informative about a method, not a clean full-precision-versus-int4 GLM-5.2 comparison.

There is also more direct GLM evidence than that old small-sample summary suggests: the grouped-int4 card reports HellaSwag normalized accuracy of 87.0% versus 83.5% for per-row int4, with 200 questions. The E8 card reports comparisons against grouped int4, but its small samples and enabled cache-aware routing do not isolate quantization error. These are useful, limited tests, not equivalence to the original model or proof of production reasoning quality.

For your own evaluation, retain the exact checkpoint, quantization, prompt template and expected outcomes. Include hard negatives, multilingual inputs and tasks that previously looped or exhausted the output budget. Check both answer correctness and completion behavior. A token-exact implementation test and a useful-business-answer test answer different questions.

Will expert streaming wear out the SSD?

Do not count 11 GB of reads as 11 GB toward a write-endurance allowance. Kingston's explanation of TBW and DWPD defines these around data written and program/erase cycles. Colibri's expert path is read-heavy, but model downloads, conversion output, persisted state and operating-system swap can still write data.

Monitor drive temperature, SMART health and swap activity during a long representative run. Keep RAM headroom so that caching does not cause write-heavy paging. A dedicated 1 TB NVMe is a reasonable planning allowance for the roughly 429 GB reference download plus operational headroom, not a software minimum or a promise that every source conversion fits. Check temporary-space requirements separately.

What is Colibri Brio mode, and when does it help?

Brio is a scoring mode for the model already loaded in Colibri, not a new model and not a Jev or Laya integration. You supply a closed set of answers. It evaluates the option tokens and returns a selected answer, relative option scores and normalized entropy rather than generating free-form text.

The endpoint supports three shapes: options for one question, questions for several questions over the same state, and schema for fields whose values must come from your lists. Here, schema is Brio's field-to-allowed-values mapping, not arbitrary JSON Schema. The server writes the JSON structure; the model only scores permitted values. A well-formed output can still contain the wrong business decision.

Zero completion tokens is not zero inference cost. The document still needs prefill and each option needs scoring. Prefix snapshots avoid some repeated processing while the engine stays alive. Restarting loses those snapshots; interleaved work and available KV slots affect reuse. The documented four-field example took 103.8 seconds with Brio versus 246.0 seconds with generation, about 2.37 times faster by arithmetic. That is one upstream Qwen3.6 experiment, not a GLM benchmark or a universal savings rate.

By default, Brio averages token log probabilities before normalizing scores across the supplied options. This reduces a simple length penalty, but does not turn the score into a calibrated probability of correctness. Low entropy means one offered option dominates; it does not prove the right option was offered. Include an abstention path, test alternative label wording and option lengths, and calibrate your review thresholds on labeled examples. Adding “human review” alone does not make the model reliably abstain.

That makes repeated classification of a shared document worth testing when local execution matters. It does not automatically beat a small classifier, a rule or a hosted decision endpoint. Our Jev decision-model review covers that separate model category; this section is specifically about using Colibri's existing engine more efficiently.

How do you install and test Colibri v1.12.0?

For a reproducible source build on a supported Unix-like system, pin the release rather than taking an unspecified future main. Install Python 3 and a supported C compiler with OpenMP first. The following does not download weights: /nvme/glm52_i4 must already contain the reference GLM-5.2 grouped-int4 checkpoint with int8 MTP linked above. Follow the release's platform instructions for Windows or a different compiler/backend.

git clone --branch v1.12.0 --depth 1 https://github.com/JustVugg/colibri.git
cd colibri/c
./setup.sh
export COLI_MODEL=/nvme/glm52_i4
./coli doctor --deep
./coli plan
./coli chat

doctor --deep checks the model's readiness; plan explains placement. Neither proves answer quality or service latency. Keep the plan and the checkpoint revision or checksums with your benchmark record. The same engine family loads GLM-5.3, but you must supply its own container and should not expect GLM-5.2's MTP behavior.

To test the gateway, start a persistent server in one terminal. Replace the placeholder with a strong local secret and keep the loopback bind while evaluating:

export COLI_API_KEY='replace-with-a-long-random-local-secret'
COLI_MODEL=/nvme/glm52_i4 ./coli serve \
  --host 127.0.0.1 --port 8000 --model-id glm-5.2-colibri

In a second terminal, set COLI_API_KEY to that same secret and submit this synthetic Brio example. The model ID must match the server. This request proposes a review queue; it does not refund a payment or establish whether a duplicate charge occurred.

curl --fail-with-body http://127.0.0.1:8000/v1/brio \
  -H "Authorization: Bearer ${COLI_API_KEY:?Set the same key as the server}" \
  -H 'Content-Type: application/json' \
  --data-binary '{
    "model": "glm-5.2-colibri",
    "state": "The ticket reports two charges for one renewal. The payment ledger has not been checked.",
    "question": "Which team should inspect the evidence?",
    "options": ["billing", "technical support", "human review"],
    "normalize": "mean"
  }'

The example is based on the documented API contract, not a live integration test performed for this article. In an application, add a timeout appropriate to measured prefill, validate the response and handle queue rejection or engine failure. Do not expose the listener publicly just to make a remote editor connect.

Can you use the API with Claude Code or a business application?

The gateway documentation describes OpenAI-compatible chat/completion routes and an Anthropic-protocol /v1/messages route. Compatibility is a protocol feature, not a guarantee of identical tool behavior or coding quality. Large agent system prompts and tool catalogs can make first-token latency much worse than a short manual chat.

The normal serving path still executes one generation at a time, with bounded admission and HTTP 429 responses for rejected or expired queue requests. Multiple KV slots preserve separate contexts where supported; they are not continuous batching or an equal feature across all engines. Size queue timeouts against measured request durations, not optimistic token-rate headlines.

Some API prose still says images and log probabilities are unsupported globally. That is too broad for the updated project: newer engines include vision paths, and Brio explicitly uses option log probabilities. Conversely, Brio does not prove that every ordinary API parameter, multimodal content block or tool feature works on every endpoint. Validate the exact engine, endpoint and input shape; unsupported combinations should fail explicitly rather than silently discard data.

Where does Colibri make business sense now?

Workload-specific evaluation, rather than a blanket production verdict
WorkloadUseful experimentAcceptance condition
Private model evaluationRun representative internal prompts without first buying a large GPU server.Useful answers and acceptable total turnaround on the exact quantization.
Repeated closed-set document questionsCompare Brio with generated JSON and a simpler baseline.Measured error rate, review workload and cold/warm latency, not only valid JSON.
Single-user local assistanceTest a smaller resident family before disk-streamed GLM.Interactive latency on that model; do not extrapolate capability from its larger peers.
Scheduled offline processingMeasure accepted items per hour, including failures and review.The queue clears inside the actual processing window.
Customer-facing concurrent APIRun load, isolation, failure-recovery and upgrade tests.Evidence for the required service level; an exposed endpoint is not that evidence.

Our assessment has changed from a blanket “not production-ready” to evaluate the workload and the selected engine separately. The project has concrete capabilities and improvements, but this review does not establish a production service level. A 25 GB machine streaming GLM-5.2 remains a poor choice for responsive, concurrent customer chat.

Local execution can improve control over model traffic. It does not by itself settle access control, retention, backups, client telemetry or permissions. For the financial decision, use our local-model versus API break-even framework rather than treating unbilled tokens as free work. Before buying hardware, our llmfit hardware guide helps shortlist ordinary model/runtime combinations; it does not simulate Colibri's expert-streaming cache.

How should you benchmark an upgrade before relying on it?

Follow the upstream reproducible benchmark protocol and retain the raw logs. Start with representative tasks, not a single greeting. Record the commit, model revision, conversion settings, command, context length, RAM/VRAM placement and storage configuration.

Run cold, warm-identical and rotating-prompt cases separately in a persistent server. Report first-token latency, whole-request time, accepted output, decode rate, cache hit rate, bytes read and peak memory. Repeat configurations in alternating order so temperature, cache history and background load do not decide the result. For Brio, also record the option set, normalization, misclassification cost and human-review rate.

Our implementation recommendation is to begin in shadow mode: compute an answer without executing the business action, compare it with an existing process, then expand only after review. The same separation between a promising demo and an operable system appears in our Twinsoft AI case study and technology-selection guide. Wavect's AI enablement service covers that implementation work, including evaluation, integration, observability and fallback design.

Sources, review date and limitations

This is a source-based technical review, updated on against release v1.12.0. Primary evidence is linked next to the relevant claims. Release-tagged code and documentation are used where available; model repositories may continue to change. We preserved the original publication date of 14 July 2026.

We did not download or run these large checkpoints, reproduce community benchmarks or test a live Brio deployment for this update. Decode-time conversions use seconds = output_tokens / tokens_per_second and omit prefill and queue time. Brio's reported ratio uses 246.0 / 103.8. Suggested acceptance criteria and business suitability are our analysis, not upstream guarantees. Where documentation disagrees, the conflict is identified rather than turned into a confident universal claim.

Frequently Asked Questions

What is Colibri for local AI?
Colibri is an Apache-2.0 inference runtime that places sparse-model weights across storage, RAM and optional VRAM. Release v1.12.0 has nine architecture-specific engine families. The C inference engine is dependency-free, but the usual launcher and HTTP gateway use Python.
Can Colibri really run GLM-5.2 with 25 GB of RAM?
Yes, but memory capacity alone does not establish usable speed. The original developer machine measured 0.05-0.1 tok/s cold. A later container-specific Core Ultra 9 / RTX 5080 run with RAM capped to 25 GB reports 0.31-0.38 tok/s. These are different historical setups, not a controlled current-release comparison.
Does GLM-5.2 need 372 GB or 429 GB of disk space?
The tagged README still says about 372 GB, but the current reference grouped-int4 file listing shows about 429 GB. Plan for the exact checkpoint revision and selected files plus headroom. The E8/IQ3 alternative lists about 289 GB and has different accuracy and decoding-cost trade-offs.
Does Colibri support GLM-5.3?
Yes. The full GLM-5.3 checkpoint uses the same engine family as GLM-5.2; its Colibri container is about 419 GB and has no MTP head. GLM-5.3-Flash is a separate engine and checkpoint. Updating the runtime does not replace existing model weights.
Will an RTX 5090 automatically make Colibri fast?
No. A GPU is useful when its supported backend reduces the measured bottleneck and enough relevant experts hit its tier. Disk misses, CPU work, memory bandwidth and long prompt prefill can still dominate. Measure the specific engine and check output quality beside throughput.
What does Colibri Brio mode do?
Brio scores a supplied closed set of answers using the already loaded model. POST /v1/brio supports one options list, multiple questions on shared context, or fields mapped to allowed values. It requires a persistent server; it is not a new model or a one-shot coli brio command.
Does zero completion tokens mean Brio inference is free?
No. Brio avoids generating free-form output, but it still processes the context and scores option tokens. Reused prefix snapshots can reduce repeated work. The documented four-field Qwen3.6 experiment took 103.8 seconds versus 246.0 seconds with generation; that result is not a universal speed or cost guarantee.
Can Brio confidence or low entropy authorize business actions?
Not on its own. The scores are relative to the options supplied, not calibrated probabilities that a business decision is correct. Low entropy can accompany a confidently wrong answer or an incomplete option set. Validate labels and thresholds on representative examples and retain policy checks and human review.
Will Colibri wear out my SSD?
Expert streaming is read-heavy, while TBW and DWPD describe write endurance. Downloads, conversion files, persisted state and swap can still write data. Monitor temperature, SMART health and swap, and preserve RAM headroom rather than treating all reads as wear-free operation.
Can Colibri run DeepSeek, Kimi and Qwen models?
Yes, where a family-specific engine exists. v1.12.0 includes DeepSeek V4 and V4.1 Flash, Kimi K3, Qwen3.6 and Qwen3.8 alongside the GLM, Inkling and OLMoE families. Support is not generic GGUF compatibility; attention, tokenizer, tensor, tool and modality support remain engine-specific.
Is Colibri production-ready?
There is no workload-independent answer established by this review. Useful local evaluation, scheduled processing or closed-set scoring should be assessed separately from a concurrent customer API. The normal serving path still runs one generation at a time. Require evidence for quality, full-request latency, isolation, recovery and cost before deployment.
Is GLM-5.2 open source?
Its published weights carry an MIT license and can be self-hosted subject to those terms. Open-weight is more precise than claiming a fully reproducible training pipeline. Colibri is the separately licensed Apache-2.0 runtime; its license does not replace the license of a checkpoint.

Final thoughts

Colibri is no longer adequately described by its first 25 GB GLM demo. Version 1.12.0 adds Brio and broader engine capabilities, while actual checkpoint sizes, backend support and benchmark methods need more careful reading than the old article provided.

The useful question is not whether a giant model starts. It is whether the exact checkpoint, hardware and workflow deliver correct, timely, operable results. Pin the version, inspect the real download, measure cold and changing workloads, and keep business actions behind explicit validation.

Need to turn these tests into a deployment decision? Discuss a measured local-AI evaluation with Wavect.

Production AI help

Building an AI product and worried about inference cost, architecture, or production readiness? Wavect helps founders turn AI prototypes into reliable production systems.

Explore the service path:

Inbox, without the noise

Follow the work that matters to you

Get a short email when we publish something new. Follow the whole blog or only the problems you care about.

What would you like to receive?
Choose your topics

Free, double opt-in, no tracking pixels.

Back
Kevin Riedl

20 min read · 14 Jul 2026
Last reviewed

Next

Get the next AI and agents field note

One concise email when we publish. No tracking pixels, and no inbox filler.

Free, double opt-in, no tracking pixels.