Magic Tools
Back to all briefs

Dev Breakfast · 2026-10-03

Today's headline: One month, 2B tokens: GLM 5.3 Flash only carried half of it. Plus 4 more: Janus: a single Go binary that runs GGUF on AMD GPUs via Vulkan; Pi 1.0 official release: ~15,000 lines of source, plus an experimental package the same day; and more.

October 3, 20269 min readDev Breakfast

Someone used GLM 5.3 Flash as their primary model for a month of coding. Of the 2B tokens, only 1B actually ran on it, and the electricity bill climbed from a planned 10 kWh to 35 kWh. When capacity runs out and you're forced to swap models on the fly — this kind of math is a genuine trap if you leave it out of your model-selection process.

🍳 Today's Headlinethe one deep dive of the day

One month, 2B tokens: GLM 5.3 Flash only carried half of it

Someone actually spent a whole year using GLM 5.3 Flash as their primary coding model — a month, to be precise. The result: of 2B tokens, only 1B ran on the target model; the other 1B leaked into other models, and energy consumption hit 35 kWh instead of the planned 10 kWh. The target model handled 50%, and the experiment itself calls the outcome a failure.

Let's start with the good news. The first half of the month did run entirely on GLM 5.3 Flash, costing $68 — about 4 kWh of electricity and 365 grams of carbon emissions, comfortably inside budget. The bad news comes in two parts. First, the "vibe coding" bill: the wrong model was picked for that experimental Wagtail MCP server, and it burned through 450M tokens, $150 and 5 kWh almost overnight. The author estimates the same result could have been had for a fifth of that cost. Second, the infrastructure couldn't keep up: GLM 5.3 Flash sits too far out on the Pareto frontier of their shortlist. Heavy demand plus capacity that can't match the big vendors hoarding GPUs, so performance degraded, and they had to fall back to DeepSeek V4.1 Flash and Qwen 3.8 Flash on the spot. Switching is easy enough — the problem is that nobody expected to have to switch.

One month, 2B tokens: GLM 5.3 Flash only carried half of it

The benchmark they incidentally released is even more worth reading: the scores of 14 models on Wagtail tasks, with DeepSeek V4.1 Flash ranking first at 95% accuracy, 14.9 Wh of energy and $0.09 per task. This is the concrete answer to the question "are cheap models good enough" — not whether they're good enough, but which one is the best value.

Put side by side, it's obvious: the cost of one wrong model pick ($150) is more than double the entire month's budget for the target model ($68). Choosing a model is already affecting the bill more than the model's per-unit price.

💡 Chef's take: Their goal for next month is refreshingly concrete — judge by cost and energy consumption, not by token count. If you're still tracking token consumption on your end, that metric probably needs replacing.

Sources:

🍲 Deep Dives · 2 more

Janus: a single Go binary that runs GGUF on AMD GPUs via Vulkan

Janus is a single-file Go binary that runs .gguf models locally and exposes an OpenAI-compatible API to the outside world. It doesn't depend on Python or Docker, and doesn't need Ollama — on Windows it's just a .exe, on Linux/macOS you get an executable via go build. Inference goes through llama.cpp's Vulkan backend, working on AMD, Intel and NVIDIA, falling back to CPU when Vulkan isn't available.

For people who write code, the interface shape is the interesting part. It listens on 127.0.0.1:8990 by default (not 8080 — the README goes out of its way to warn about this), /v1/chat/completions and /v1/models are directly compatible, and with Base URL set to http://127.0.0.1:8990/v1 and the API Key left empty, you can plug straight into Cursor / Cline. /models/load supports hot-swapping models without restarting the process; the model path is set via JANUS_MODEL_PATH, JANUS_GPU_LAYERS=-1 puts all layers on the GPU, JANUS_VRAM_CEILING_MB defaults to 9216, and JANUS_MAX_TOKENS defaults to 4096. Reasoning models with `` get their chain of thought split out into reasoning_content, and the chat template is read automatically from the GGUF metadata. Each model is usually 2–8 GB on disk; Janus plus llama.dll is about 50 MB.

The gotchas are stated just as plainly: on Windows, Go can't overwrite a running .exe, so if a code change doesn't take effect, go kill all janus.exe processes in Task Manager first — that's also the cause of occupied ports. The first conversation takes 10–60 seconds; that's the model loading into VRAM, and Q4 quantization makes it faster. On Linux, libllama.so needs to sit next to the binary or be on LD_LIBRARY_PATH. The original post says nothing about how it behaves with multiple GPUs or large models, and gives no performance numbers at all — it's a shell wrapped around llama.cpp, its ceiling is llama.cpp's ceiling, so don't expect it to be faster than upstream.

Sources:

Pi 1.0 official release: ~15,000 lines of source, plus an experimental package the same day

Earendil has pushed Pi to 1.0. The official line is that hundreds of thousands of people use it every week, and what this stable release brings in includes: Codemode with native MCP support, virtual model extensions, lazy-loaded tools, cache warming for Anthropic models, system messages inserted mid-conversation, new TUI themes, and fullscreen by default. Installation is still a one-liner curl or PowerShell command, MIT licensed, docs at pi.dev.

What's interesting is their sense of rhythm: agentic tools change every week, but most of those changes don't survive — so Pi chooses to wait until something is proven before considering bringing it in. Pi Durable, released the same day, is the product of exactly that philosophy. It isn't a replacement for Pi, but a framework for long-running, persistable agent applications: the storage backend ships with memory, SQLite and JSONL; the SQLite and JSONL parts don't depend on the Node API and can run on Bun or Cloudflare Durable Objects with a small adapter; every execution step is a task with a checkpoint, so if the process dies, a new process opening the same store resumes from where it left off — truncated model requests get resent, truncated tool calls with side effects safely re-run, and requestId guarantees exactly-once commit. The whole source, excluding tests, is about 15,000 lines — roughly 150K tokens for GPT and 250K for Claude. That sounds large, but the official line is that the 3,000 lines of storage backend are usually skippable; when you're actually having an agent read its own source and modify itself, you don't need to feed it all of it.

For people who write code, the thing worth noting is that it clarifies what "harness" means: storage + a mechanism for running multiple conversation turns in parallel + tools + an execution environment. The four-piece set. If you want your agent to run across machines, or want multiple people steering the same agent at once, this abstraction is directly reusable. As for whether to adopt Pi Durable right now, my suggestion is to first work out whether your scenario is genuinely "long-running". If you're just running a coding agent in the terminal, 1.0 has already stabilized what needed stabilizing, and there's no reason to swap out your dependencies for a new package. It's an experimental package — the name says experimental, so don't blame anyone but yourself when something breaks.

Sources:

🥢 Sides · 2 more

Bez wants to auto-generate a browser engine from specs; current coverage is 0.6%

The project is called Bez, and the idea is to generate a Web rendering engine from spec text and WPT tests, cross-validating against three shipped browsers, with Rust accounting for 76%. The goal it sets is refreshingly concrete: the default build is a full engine, and you can also trim it down to the features a site actually uses — code that never runs simply doesn't make it into the binary, so the engine is small, light on memory, easy to embed, and gets regenerated when the spec changes instead of being patched by hand. But looking at the status table it publishes itself, against the 17,259 leaf keys in browser-compat-data 8.0.4, the generated part covers only 0.6%, the hand-written part 0.3%, and 93% is untouched — with HTML, JavaScript, SVG and WebAssembly all 100% uncovered. Hand-writing an engine takes hundreds of engineers for years; I'll grant that judgment. But between 0.6% and running a real page lies not a single pipeline, but a decade's worth of accumulated compatibility mud. First understand how it turns a spec into verifiable code, then decide whether it's worth following.

Sources:

Four code smells of agentic coding: the code goes rotten, people go cold

Agentic coding works well, but the author lists four failure modes, and the first one stings the most: code produced by LLMs carries a "slop" smell. Not strictly worse — visibly different — and that sloppiness doesn't look like growing pains; it looks like the model's signature move. The consequence is that the codebase becomes repellent to humans: once an agent is let in, it quickly turns a space once shared by people into an AI wasteland. The second is the growing distance between engineers and their code: no longer written by hand, no longer read — you just issue instructions and skim a report, so you stop feeling ownership of the output, features get half-thought-out, bugs get Band-Aids, and you leave work feeling hollow. The original post gives no numbers and no solutions. I agree with the direction of this observation, but I'd rather know: which parts of the process should have "a human must go through it by hand" written into them, instead of vaguely mourning the feel of doing it yourself.

Sources:


If tomorrow you had to bet your entire primary model setup on one vendor, would you bet it holds up, or would you keep a fallback ready to switch at any moment? See you tomorrow at 8.

This edition selects 5 items from the past 24 hours across X / Hacker News / GitHub Trending out of 59 total (hourly reporting throughout the day, fact-checked, then curated in the morning). Content is LLM-assisted and each item comes with a link to its original source; please cross-verify for important decisions.

Like this brief? Get it by email

Daily AI coding picks at 8:00, plus a hands-on field-notes issue every Saturday. Written in Chinese.

This page is auto-generated by LLM aggregation; please cross-check with original sources.

Dev Breakfast · One month, 2B tokens: GLM 5.3 Flash only carried half of it | Magic Tools | Magic Tools