Magic Tools
Back to all briefs

Dev Breakfast · 2026-10-10

Today's headline: LiteLLM Swaps in a Rust Core: One Gateway to 100+ LLMs, and How the 8ms P95 Was Calculated. Plus 4 more: 1500 Lines Split Into 10 Subtasks, and the Agent Still Runs Them One at a Time; One Swift Binary: Let Your Agent Draw Arrows on the macOS Screen; and more.

October 10, 202610 min readDev Breakfast

LiteLLM has swapped its core to Rust, while the Python layer — import litellm and completion(model="openai/gpt-4o", ...) — doesn't need a single line changed; what's running underneath is already a different beast. When I saw that "8ms P95, 1k RPS" figure, my first instinct was to go click the benchmarks link — the gateway computing fast on its own doesn't mean your call from Shanghai to some provider is fast either.

🍳 Today's Headlinethe one deep dive of the day

LiteLLM Swaps in a Rust Core: One Gateway to 100+ LLMs, and How the 8ms P95 Was Calculated

LiteLLM's README now reads "The fastest, litest AI Gateway. Rust core with Python SDK." The core is Rust, but the outside is still that same Python API via import litellm — the completion(model="openai/gpt-4o", ...) you already wrote needs not a single line changed, while what runs underneath is no longer the same thing. The problem it solves hasn't changed: 100+ LLM providers, each with a different SDK, authentication, request format, and error types; LiteLLM unifies them into the OpenAI format. You can embed it as a Python SDK directly in your code, or deploy the Proxy Server as a team-level AI Gateway — virtual keys, spend tracking, guardrails, load balancing, logging, and an admin dashboard all come out of the box.

The number worth watching most closely is that "8ms P95 latency at 1k RPS" — a benchmarks link hangs right next to it, but latency always depends on what hardware you're on, which provider you're proxying to, and which network hop you're in. The gateway computing fast on its own doesn't mean your call from Shanghai to a model on Bedrock takes only 8ms. The number only means something if it matches your scenario; if it doesn't, it's just marketing copy.

LiteLLM Swaps in a Rust Core: One Gateway to 100+ LLMs, and How the 8ms P95 Was Calculated

What actually affects you is the change in distribution. A separate litellm-core distribution now appears in the README: same import litellm API, same runtime dependencies, but no optional extras, no CLI entry point, no bundled dashboard. In the same environment you can install only one of litellm and litellm-core, because the two own overlapping Python files. If you want the proxy, CLI, and optional dependencies, use litellm; if you only want the SDK, install core. And the release integration for core is still in flight — for now you have to build it from the repo yourself: python scripts/build_core_distribution.py --out-dir dist/core, then pip install dist/core/litellm_core-*.whl. The build machine needs Git, uv, and a Rust toolchain — meaning a pure Python project now requires Rust to produce a package from source. The version number is read from the root pyproject.toml; the build doesn't modify source files and produces a wheel plus a self-contained sdist.

Moving gateway-type middleware from Python to Rust makes sense on paper: it's the mandatory path for every request, every extra hop adds latency, and trading a compiled language for throughput is standard practice. But the costs are standard too — a longer build chain, more complex binary distribution, and when things go wrong, the stack you can read changes from Python to Rust. Swapping the core: the gains live in the load test report, the costs show up the next time your CI fails.

💡 Chef's take: The release integration for that litellm-core is marked "pending" in the README, so right now installing core means running the build script yourself; the manual step only disappears once it's officially published to PyPI.

Sources:

🍲 Deep Dives · 2 more

1500 Lines Split Into 10 Subtasks, and the Agent Still Runs Them One at a Time

Someone wrote a long post cataloging the failings of coding agents, and it kicked up a lively discussion on Hacker News. The author's core claim: models keep improving, agents haven't kept up, and the bottleneck is the agent layer. His example is very specific — adding a "password-protect links" feature to an open-source file-sharing app, about 1.5k lines of new code in total; OpenCode obediently split the task into 10 subtasks and then dutifully did them one by one.

That's the most jarring part: those 10 subtasks are independent of each other, a textbook embarrassingly parallel workload — a single machine can spin up dozens of processes at once, context switching is millions of times faster than a human's — and yet it queued them up. Claude Code fares a little better, spinning up one or two subagents, but it also waits for them all to finish before moving on; the author says several times a day he watches it sit idle waiting for an end-to-end test run, and only once the test passes does it remember "oh, I should look at the git history to learn your commit message conventions."

Two other problems deserve more of your attention. First, it can't delegate: when a task needs scanning 50k lines of code for a pattern, the agent won't say "this job just needs a cheap fast model," nor "this one's too hard, swap in a smarter one" — it'll just grind away with the most expensive model, even though 95% of the work is grunt work. Second, it doesn't know itself: ask Claude how to use a Claude feature and it has to search the web, without even checking the version number — yet it will silently download a 13 GB file for a feature you've never used without batting an eye. Saving 50 KB of docs at install time costs you a fresh search every single time.

For people who use these tools daily, the conclusion isn't "don't use them" — it's to stop expecting them to manage tasks for you. Splitting, ordering, and model selection all still have to be carried on your own shoulders for the near term; writing clear prompts, cutting tasks small enough, and manually swapping in cheaper models on the right subtasks are three moves that save real time and real money right now. Agents will get better, but until they learn to delegate on their own, keep your hands on the wheel.

Sources:

One Swift Binary: Let Your Agent Draw Arrows on the macOS Screen

An Agent can refactor a monorepo, write database migrations, and explain monads to you clearly — but the moment it needs you to click an "Allow" button, all it can do is print "please click Allow in the dialog" into a terminal you're not even looking at. That gap is exactly the one centimeter bigarrow wants to fill.

It's a macOS command-line tool, plus a skill for Claude Code and Codex, and it does three core things: draw arrows over any window, draw boxes, and write big text. Clicks pass through, keyboard focus doesn't move, and when the job is done the arrow disappears on its own. MIT licensed, a single Swift binary, no daemon, no menu bar icon, no login, no telemetry — the author even double-checked twice to confirm there's no AI in it; it's just an arrow.

The practical payoff is here: drawing things requires no macOS permissions at all. It uses a transparent window layered above every display and every Space, so you never have to go into System Settings to grant Accessibility access — and that's a key point, since the permission prompt itself is often the very problem it's trying to solve for you. Element targeting is done via --element plus --app, with --role (button, radiobutton, etc.) also available; styling includes --style ring/box, --shape zigzag, --color, --size; direction via --from right/bottom-right/top; it can even attach a close button. The README examples are refreshingly real: six arrows teaching you how to turn off "Show Desktop by Clicking Wallpaper," a green arrow pointing at Terminal's toggle, a three-step walkthrough of Keynote's animation flow, and a pink arrow helping Mom find "Save as PDF" in the print dialog — with a small black arrow beside it pointing at "Cancel," labeled Not this one, Mom.

For people who write code, the value of this thing isn't the arrows themselves — it's that it makes "the most awkward stretch of human-computer collaboration" programmable. You're running a pipeline that needs human confirmation, the agent is stuck on an authorization dialog, and where you previously could only refresh logs in a terminal, now it can point straight at the thing for you. The costs are clear too: it's currently macOS-only, and this transparent-window-plus-click-through trick would have to be rewritten from scratch on Linux and Windows; element targeting relies on accessibility information, so any UI redesign can leave it pointing at thin air. So don't expect it in production workflows — it's better suited as a pointing finger during local development. After all, teaching an Agent to point is much safer than teaching it to click for you.

Sources:

🥢 Sides · 2 more

Microsoft Open-Sources MXC: A Sandbox for AI-Generated Code

Would you dare run code that AI wrote for you, straight up? Microsoft has open-sourced MXC, a cross-platform sandbox execution system for Windows, Linux, and macOS, built specifically to run untrusted code — model outputs, plugins, and tools all qualify. It offers multiple isolation backends including ProcessContainer, Windows Sandbox, LXC, Bubblewrap, Seatbelt, and MicroVM, configured via JSON config files and network policies, and ships with three SDKs: Rust, .NET, and Node. Windows Sandbox, MicroVM, and Hyperlight are marked experimental. For those of you letting Agents execute code every day, this tackles exactly the "lock it in the cage before running it" problem. But a sandbox isn't a universal lock — misconfigured policies leak just the same, so think through what needs isolating before choosing a backend.

Sources:

Anthropic Will Publish Claude Behavior Reports Regularly: First Report Lists Four Categories

Anthropic says it will start publishing model behavior reports more frequently, reaching beyond system cards and routine risk reports. Today's first edition describes four categories of behavior they identified in evaluations and internal use, in each of which Claude "acted on real" — the original text cuts off right there; which four categories and what the consequences are, left unelaborated.

What's worth noting here is the posture of the publisher: shifting from managing how users use the model to proactively disclosing what the model itself will do. A model vendor willing to put its own model's behavior issues on the table regularly carries more information than publishing capability leaderboards alone. But with only a bare opening given for the four categories, judgment has to wait for the full report. My consistent view: a model's behavioral ceiling is determined by data, not by marketing, and the value of this kind of report lies entirely in the details.

Sources:


With the gateway layer moving to Rust, will you upgrade the LiteLLM setup you have in place in one go, or pilot it on a non-critical service first? See you tomorrow at 8am.

This issue selected 5 items out of 53 collected over the past 24 hours from X / Hacker News / GitHub Trending (sampled and written hourly throughout the day, fact-checked, then compiled in the morning). Content is LLM-assisted; every item links to its original source, so cross-verify for important decisions.

Like this brief? Get it by email

Daily AI coding picks at 8:00, plus a hands-on field-notes issue every Saturday. Written in Chinese.

This page is auto-generated by LLM aggregation; please cross-check with original sources.