Dev Breakfast · 2026-10-05
Today's headline: macOS 27 kills the Apple Intelligence master switch: one script helps you evict the models. Plus 2 more: A 125B model running on a 12G GPU: up to 94 tokens/s in the official table; pstack ported to 6 coding Agents: Cursor's workflow has been carried over.
macOS 27 has killed the master switch for Apple Intelligence. The features are off, yet the base model, Genmoji, and the models behind Xcode completions are still squatting on disk and refusing to leave. RemoveMacAI goes through Apple's own configuration profile plus download redirection — not by hard-deleting system files — and that kind of front-door bypass beats a one-click cleaner.
macOS 27 kills the Apple Intelligence master switch: one script helps you evict the models
In macOS 27, Apple Intelligence no longer has a "one-click off," and once you turn the features off, the models already downloaded locally still sit on disk. Someone built RemoveMacAI — a single command turns the features off, deletes the models, and seals off macOS's ability to re-download them. What it removes is specific: the Apple Intelligence base model, plus the individual models for image generation, Genmoji, Spatial Photos, Photos Clean Up, and Xcode code completion. The features it turns off include Siri (including "Hey Siri" and the menu bar icon), Writing Tools, Genmoji, Image Playground, ChatGPT extension, summaries in Mail/Messages/Safari/Notes/Notifications, Smart Reply in Mail, inline text prediction, Spatial Photos, Photos Clean Up, and Xcode predictive code completion.
Its approach is not to hard-delete system files, but to go through Apple's own configuration profiles and the asset service: a profile applies Apple's restriction keys for Apple Intelligence, and then redirects the download URL of each removed model to a closed local port, so macOS won't quietly pull it down again. SIP stays on, not a single file under /System is touched, and removing the profile restores everything — the system will simply re-download a model whenever a feature needs it. Installation is a one-line curl or Homebrew; the script verifies SHA-256, runs in a temporary directory, and leaves no install trail. removemacai status lists the space taken by each feature and model, --keep preserves specified features, --dry-run only previews, and revert rolls everything back in one step. Apple silicon is required, support is macOS 27 only — 27.0 tested working, 27.0.1 needs 0.2.3 or newer.

Two things are worth clearing up first, so you don't install it and wonder "why hasn't anything changed?" First, Storage settings may still show Apple Intelligence's usage: the asset service frees the models immediately, but macOS deletes files on its own schedule, and until then "System Settings > General > Storage" still counts it. Second, don't assume the job is incomplete just because you see a process called Siri in the process list — in macOS 27 the Spotlight window itself runs under a process named Siri, and some system services are SIP-protected and could never be uninstalled anyway. The cost is stated plainly: the features listed above will stop working, and so will any app relying on Apple's on-device models (the Foundation Models framework, the "Use Models" action in Shortcuts), Visual Intelligence, and natural-language editing in Calendar. Voice dictation is unaffected — that lives in a separate setting — and the voice models weren't deleted either.
This project solves "can't be turned off" and "can't be fully deleted," but how much space it saves you has to be seen by running status yourself — the README gives no numbers. If you want it clean but don't trust a third-party script, you can write the profile yourself following the same approach; if you want convenience, run --dry-run for a look before deciding.
💡 Chef's take: The thing worth examining in tools like this is their removal path. It goes through Apple's asset service instead of
rm, so system updates won't wipe out your changes, and the profile plus its download blockade survives into the next major version.
Sources:
A 125B model running on a 12G GPU: up to 94 tokens/s in the official table
Strata, an open-source inference engine, has squeezed Qwen3.8-Flash-Next onto consumer GPUs. Officially, it needs 12 GB of VRAM as a minimum and 32 GB of RAM, installs in one click on both Windows and Linux, and can spin up an OpenAI/Anthropic-compatible endpoint locally. The two machines tested officially are an RTX 5070 (12 GB) and an RX 9070 XT (16 GB); the smaller the quantization level, the faster it runs — Q2_0 writes answers at 94 tokens/s on the 5070 and reads a 32K prompt at 2650 tokens/s, while IQ3_S drops to 53 tokens/s. The AMD machine at the same quantization only manages 60 tokens/s. The original post also says outright that cards with more VRAM are faster, and a 24 GB 3090 can reach roughly 100–140 tokens/s. As for the claim of "125B at 100 tokens/s," no tier in the official table hits 100 tokens/s — that is the estimated range for the 3090, so don't treat it as measured. If you want to run locally without paying an API bill, it's worth understanding the speed difference the quantization tiers buy you.
Sources:
pstack ported to 6 coding Agents: Cursor's workflow has been carried over
In earlier issues we talked about Offrun trying to manage all coding Agents in one place, and about Cloudflare's Agent workbench — both of those took the "unified workspace" route, first consolidating containers and entry points. Today's story is another concretization in the same direction: someone has ported Poteto's pstack into six versions — Claude Code, Codex, Pi, OpenCode, Gemini, and Prime Agent — translating Cursor's primitives onto each vendor's harness. The difference is that Offrun manages "where it runs" while pstack manages "how it runs": the same rigorous Agent workflow, so you don't have to relearn it when switching Agents. The repo gained 242 stars on GitHub trending in a single day, bringing its total to 1.1k.
The word worth pondering is "translating." The fact that Cursor's primitives can be carried over shows that the workflow layer is being decoupled from the harness; but translation always comes with loss. Each Agent's tool calling, context management, and permission model differ, so the same flow won't feel the same across different harnesses. If you want to try it, pick the one you use most day to day rather than rolling out all six at once.
Sources:
Would you run this command on your own machine and evict the models along with Siri, or would you rather keep the master switch's stand-in around? See you tomorrow at 8.
This issue selected 3 items from 39 pieces of information across X / Hacker News / GitHub Trending over the past 24h (sampled and written hour by hour throughout the day, fact-checked, then compiled in the morning). Content is LLM-assisted, each item comes with a link to its original source, and important decisions should be cross-verified.
Like this brief? Get it by email
Daily AI coding picks at 8:00, plus a hands-on field-notes issue every Saturday. Written in Chinese.
This page is auto-generated by LLM aggregation; please cross-check with original sources.