← Founder Notes
Archive

Colibri, a tiny c engine to run frontier mixture of experts models on a normal machine, topped…

Yethikrishna ROriginal on Threads

colibri, a tiny c engine to run frontier mixture of experts models on a normal machine, topped github trending on september 14 with no cloud and no api. the largest models are being shrink-wrapped into single-file runtimes.

local moe inference just became a hobbyist project.

Context

The Colibri README, read on 4 October 2026, describes a pure C engine with zero engine dependencies that streams experts from disk, with one C file per model family and seven families running: GLM-5.2 at 744B, GLM-5.3-Flash, Inkling, Kimi K3, DeepSeek V4 Flash, Qwen3.6 and OLMoE. It lists the Apache 2.0 license. Its author-reported speeds include 4 tokens a second for the 744B model on a stated configuration, 5.8 to 6.8 on six RTX 5090s, about 1.8 on a 128 GB CPU-only desktop, 1.07 on a single RTX 5070 Ti and 0.05 to 0.1 on a 25 GB box, and it states no SLA on speed. Gigazine of 10 July 2026 reports it running GLM-5.2 on a 25 GB PC.

How it compares

The speeds are the author's, tied to specific configurations, and not comparable across rows. Normal machine holds for some models at very low speeds, and frontier here means large open-weight mixture-of-experts models, not a hosted-model comparison. It is one C file per family, not one file overall. The project existed by 10 July 2026, and no dated evidence for topping GitHub trending on 14 September was found, so that is unverified. Local mixture-of-experts inference becoming a hobbyist project is the author's opinion.

Watch next

  • A dated GitHub Trending capture and independent tokens-per-second tests.

Sources

  1. Colibri README (GitHub)github.com
  2. Colibri runs GLM-5.2 (Gigazine, 10 Jul 2026)gigazine.net

Provenance

The note above is reproduced unedited from the original post, first published on Threads on 19 September 2026 at 04:19 IST. Sources are the papers and datasets the note draws on.

View the original post
Embed this note
<iframe src="https://founder.myndlabs.tech/notes/embed/colibri-a-tiny-c-engine-to-run-frontier-DdcjA6VCJAB" width="480" height="420" style="border:0;max-width:100%" loading="lazy" title="Colibri, a tiny c engine to run frontier mixture of experts models on a normal machine, topped…"></iframe>

More notes