Dev Breakfast · 2026-10-04
Today's headline: Remove "think carefully" from your CLAUDE.md: Opus 5.5's official tips list. Plus 4 more: Cloudflare open-sources its Agent workbench: every app runs only for you; Gemini kills free usage for Flash and Pro; and more.
The most counterintuitive tip in the official Opus 5.5 guide: delete every "think carefully" and "think step by step" from your prompts and CLAUDE.md, because the model decides for itself how long to think before each reply — nudging it only slows it down. My take: this isn't just about saving a few seconds — it shows that some of the prompting habits we've accumulated over the past two years have quietly turned into pure ritual.
Remove "think carefully" from your CLAUDE.md: Opus 5.5's official tips list
On September 22, Opus 5.5 shipped an official usage guide covering how to use the same model in the Claude app and in Claude Code. The most counterintuitive piece: the official advice is to delete all "think carefully" and "think step by step" from your prompts and CLAUDE.md — Opus 5.5 already decides for itself how long to think before each reply, and asking it to "think carefully" only makes responses slower. The official test conclusion in the chat product: delete that line and replies start sooner, with no measurable drop in quality.
Worth paying even more attention to is the "stopping point" issue. Opus 5.5 will proactively report progress on long tasks, but sometimes it stops and asks whether you want to continue — a summary that only reports the next step without executing it, an invitation to continue, a pile of options that actually don't block any work. The official fix is to hard-code the rule in CLAUDE.md: keep running without waiting for my input, and put the status explanation and the next action in the same message; only stop and ask when you cannot proceed independently on your own, or right before destructive operations like deleting data, force-pushing, or modifying files outside the repository. The official guidance also recommends keeping permission confirmations for destructive commands.

Three takeaways you can copy directly. First, hand over the entire task in one go and spell out what "done" looks like — for example "every endpoint uses the new client, the old client is removed, all tests pass" — and when it should stop and ask you. Early testers let it run multi-hour coding tasks with almost no supervision. Second, for large-scale audits or migrations, have it assign each service to a separate subagent, verify the evidence returned by each subagent before adopting it, and finally consolidate everything into a single table. Third, on long tasks, have it write the checklist into TASKS.md and tick items off one by one — long runs fill up the context window, after which Claude Code summarizes the older turns, and a checklist written to a file survives; you read the file instead of scrolling through the transcript. Also, if you think of something mid-run, you can just keep typing and hit enter to append it — no restart needed, because each run now lasts longer and the cost of restarting is higher.
There's one detail worth recording separately for design-type requests: saying "don't make it too generic" is basically useless — it just swaps one default style for another. Instead, list the specific patterns you don't want to see: no off-white backgrounds, no italic emphasis words in headings, no "01 / 02 / 03" numbering, no monospace labels, no pill-shaped buttons.
The most valuable line in the official guide is "delete think carefully"; the most easily overlooked is "write clearly in CLAUDE.md when it should stop and ask you."
💡 Chef's take: What's easiest to miss in this guide is how it draws the line around "destructive operations" — deleting data, force-pushing, touching files outside the repository. For these three categories it recommends keeping permission confirmations; don't tear down that gate just to save a few popup clicks.
Sources:
Cloudflare open-sources its Agent workbench: every app runs only for you
Cloudflare has open-sourced the AI workbench it has used internally for years. The repo is called cloudflare-os, and it's positioned as an "AI productivity operating system." According to the README, more than half of Cloudflare's people, from engineering to sales, use it daily; the current version is v2, a complete rewrite released in August 2026, and the official description is "early access" — the capabilities are already decent, it admits, but the edges are rough.
What's really worth looking at isn't the chat box but its redefinition of "applications." When you build a deck in it, the system doesn't call out to some SaaS — instead it generates a private instance of a presentation app just for you, running in an isolated sandbox, which the official name calls a Gadget. There are two benefits: if someone else's presentation app has a security hole, it can't reach you; and if you want to add a feature, you just have the Agent modify the code of your own instance — and if it breaks, only your copy breaks. The security layer is called Gatekeepers, essentially a beefed-up MCP server responsible for OAuth, narrowing permissions down to the single resource the user specified, and logging every call. One design in there strikes me as more practical than the marketing: for operations that need human confirmation, the Gatekeeper first simulates the result locally so the Agent can keep going, and you batch-approve or batch-reject them when it's convenient. This solves the classic scene of "dispatch a task to the Agent, go pour a coffee, come back and find it stuck at step one waiting for approval" — which is why everyone just runs --dangerously-skip-permissions and pretends not to notice. Getting started is cheap: install and run pnpm run-local, and you can see the whole picture on port 8787 locally, or deploy straight to your own Cloudflare account.
But the word "operating system" is used too loosely. It's neither a traditional OS nor a product you can deploy in your company and be done with it — the official intent is for you to fork it and turn it into "your company's OS," and you bear the cost of that transformation yourself. So don't rush to move your internal toolchain onto it; run it locally first and focus on whether the Gatekeepers' permission model and the sandbox boundaries are what you want. If you do want to connect it to your own systems, writing Gatekeepers is work you probably can't avoid. For people who write code, it reads more like a reference implementation you can take apart — how an Agent is boxed into a sandbox, and how side-effecting operations get asynchronous approval — those two pieces are more valuable than the UI.
Sources:
Gemini kills free usage for Flash and Pro
Someone on Reddit posted a message saying Gemini is ending free use of the Flash and Pro models. The original text was just that one line — no effective date, and no indication of whether the free tier is being shut down entirely or merely narrowed. For people who write code, the weight of this depends on whether you've pushed Gemini into your own toolchain — if you call it from scripts for summarization, code completion, or batch translation, go check your API bill and call volume first, rather than waiting for a day when requests start failing outright. Free credits have never been a promise; they're an acquisition budget, and when the budget runs out they get pulled. I've always planned around "it could vanish any day."
Sources:
System76 says no to AI-generated code across most of its COSMIC codebases
System76 has set a rule for the COSMIC desktop environment that ships with Pop!_OS: most of its codebases will not accept AI-generated code contributions. The original post only contained a title and a one-line summary, with no detail on how it's defined or how it's verified. For people who write code, this has two layers: first, if you're submitting a PR to this kind of project, you need to confirm whether the autocomplete and generation tools you use count as banned; second, it puts the question of "who's responsible for AI-written code" on the table. Evaluating programmers on AI generation rate is itself a way of treating a tool as a grade, and I'm against that kind of metric.
Sources:
Greg KH on kernel security: what to defend against in the LLM age
This is a talk by Greg Kroah-Hartman titled "Security in the LLM Age." The original page only captured YouTube's verification page, with no body text, so what was actually covered and which examples were given can't be confirmed. Two things can be established: the speaker is a long-time maintainer of the kernel stable branch, and the topic is LLMs and security. For people who run kernels and watch CVEs every day, that combination is worth clicking through — security discussions in kernel land usually lag the noise in application-layer land by half a beat, but they land more solidly when it comes to patches. As for the conclusions, wait until you've watched before deciding; don't take a title as an opinion.
Sources:
The official word is that deleting "think carefully" doesn't measurably hurt quality — do you buy it? I'm at least planning to delete one line from my own CLAUDE.md and find out. See you tomorrow at 8.
This issue picked 5 items out of 44 pieces of information from X / Hacker News / GitHub Trending over the past 24h (collected and written hourly throughout the day, fact-checked, then compiled in the morning). The content is LLM-assisted, with an original source link attached to each item; please cross-verify for important decisions.
Like this brief? Get it by email
Daily AI coding picks at 8:00, plus a hands-on field-notes issue every Saturday. Written in Chinese.
This page is auto-generated by LLM aggregation; please cross-check with original sources.