Magic Tools
Back to all briefs

Dev Breakfast · 2026-09-29

Today's headline: Adding 'Do not guess' cuts hallucination rate from 71% to 20%. Plus 4 more: Go's import path tied to GitHub: how much code changes when switching hosting; Sonnet 5.5 released: Terminal-Bench jumps from 10.3% to 70.6%; and more.

September 29, 20267 min readDev Breakfast

Adding a single phrase 'Do not guess' to the prompt, with the same batch of models and the same task, cuts the hallucinated claims from 71% to 20%. I believe in this direction: many so-called hallucinations are because we didn't ask the model to shut up; with proper constraints, the model is willing to leave blanks for what it doesn't know.

🍳 Today's Headlinethe one deep dive of the day

Adding 'Do not guess' cuts hallucination rate from 71% to 20%

Someone tried a trick in a benchmark: adding 'Do not guess' to the prompt reduced the model's hallucinated claims from 71% to 20%. Same batch of models, same task, only changed this sentence. On the other hand, Anthropic's Opus 5.5 prompt documentation also discusses the same thing—output tokens are over 30% faster than Opus 5, but behavioral differences need to be calibrated with prompts and harnesses, such as marking user-pasted text or adding progress signals for unattended agents. The model's upper limit is determined by data, not promotion; a single constraint can cut more than half of the fabrication, indicating that many 'hallucinations' are simply not being asked to shut up.

Adding 'Do not guess' cuts hallucination rate from 71% to 20%

Sources:

🍲 Deep Dives · 2 more

Go's import path tied to GitHub: how much code changes when switching hosting

Go has a rather likable design: the import path itself is the download address for the code. When you write import "github.com/thetrueares/boneclone", the Go toolchain knows to fetch from GitHub. The advantage is that finding bug reports and distributing libraries doesn't require centralized package management; the downside is—your code is now tied to the hosting provider.

Author Iain Cambridge explains this cost bluntly: if you want to move from GitHub to GitLab one day, you must change the code, otherwise you'll still fetch the old version. Once the migration work piles up, people just skip it. He has seen a company using GitLab, GitHub, and Azure DevOps simultaneously, not because it's technically necessary, but because changing code locations is too much work and they 'don't have time', so they'd rather pay extra hosting fees. This is no longer an aesthetic issue; it's a billing issue.

The solution is to assign a custom domain to your package, like go.iain.rocks, go.uber.org, go.mongodb.org, etc. Users don't need to worry where the domain points—go.iain.rocks/boneclone now points to the GitHub repository, and when switching to GitLab one day, the installation command doesn't change at all. Implementation relies on two things: an Nginx configuration that checks if the request has the go-get=1 query parameter; if yes, it passes to the Go toolchain to read the HTML; if not, it redirects human visitors with a 301 to GitHub; plus an index.html that uses go-import and go-source meta tags to declare 'where the repository for this domain is and how to construct the source browsing link'. The author posted his configuration as-is, so you can use it by changing the domain.

His conclusion is: every team doing commercial development with Go should use custom domains for internal libraries and packages. I agree with this, and you don't need to think big—if you don't migrate now, it doesn't mean you won't migrate later, and the cost of this increases with the number of references. Secure the domain first, and the initiative remains in your hands.

Sources:

Sonnet 5.5 released: Terminal-Bench jumps from 10.3% to 70.6%

Anthropic has released Sonnet 5.5, the second model in the Claude 5.5 family. The official stance is 'over 30% faster than Sonnet 5 and up to 30% lower cost for most tasks'. But the striking numbers are not in the marketing copy but in the evaluation tables: in Terminal-Bench 4.0, an agentic coding evaluation, Sonnet 5 only scores 10.3%, while Sonnet 5.5 jumps to 70.6%. The gap within the same model generation is nearly sevenfold, which doesn't seem like a 'clear upgrade' but more like finishing work left undone from the previous version.

The price hasn't changed: $2 per million input tokens, $10 per million output tokens, and $0.20 per million cached reads, same as Sonnet 5. The so-called '30% cost reduction' is not a price cut but fewer tokens used for the same task—the official wording is 'typically needs far fewer tokens', saving up to 30% per task in tests. This distinction is worth remembering: bills are calculated by tokens, the unit price hasn't changed, and the savings depend on how many tokens your tasks actually save, not a blanket 30% discount. Feedback from Slack indicates about 14% fewer output tokens, which can serve as a reference anchor. Moreover, its positioning differs significantly from Opus 5.5: Sonnet 5.5 excels at clear-boundary daily tasks, bug fixing, documentation, and spreadsheets, while Opus 5.5 remains significantly stronger in open-ended tasks requiring continuous judgment. The official documentation itself notes below that 'benchmarks only reflect one facet of capability'.

There are two details more important than benchmark scores. First, the default effort level: Claude Code and App default to Medium, Platform defaults to High. In the official charts, Sonnet 5.5 at Low or Medium surpasses Sonnet 5's best score, and the cost per task is less than one-tenth of the latter—meaning, even without changes, you might already be saving money, but if you manually max out the level, some of those savings will be returned. Second, on the security side: this is the first Sonnet model launched with cyber safeguards, given that its cybersecurity capabilities are comparable to Opus 5. The official statement is that such protections target only a small subset of high-risk requests; daily development and most life sciences work are unaffected—put belief aside, if your pipeline has security-related automation tasks, it's worth running a test to see if anything gets blocked. Haiku 5.5 is officially said to join in a few weeks, so those doing high-concurrency, cost-sensitive scenarios might wait before choosing.

Sources:

🥢 Sides · 2 more

Imp: Moving DSPy to BEAM, where an Agent becomes a process

The approach of DSPy, which 'writes each model call as a typed, measurable function', now has an Elixir version. Imp is a full port: signatures, modules, optimizer, agent loop, retrieval are all there, running on OTP. The difference is in the Agent's form—within BEAM, it's just a process, holding its own state, receiving messages, and supervised under a supervisor, managed alongside your application. The way of writing code also changes: no prompts, no parsers, using Imp.signature to declare inputs and outputs, and fields like kind are either enum values or directly throw errors. The optimization step offers several paths like GEPA, MIPROv2, SIMBA, etc.; GEPA reads failed samples to rewrite instructions, and max_metric_calls: 300 is its budget limit. Modified programs can be saved as JSON and reviewed as diffs. To decide whether to adopt it, look at two things: whether your tech stack is already on BEAM, and whether you're willing to add an extra layer of declaration for 'measurability'—the value of the ported version isn't in feature alignment, but in gaining concurrency and fault tolerance for free.

Sources:

Everyone asks Claude: code runs, nobody knows why

An 871-word short article cites a large company engineer's self-narrative: after half a month on the job, specs, code, tests, PRDs, tickets, reports are all generated by Claude Code, from L1 to L7 doing the same thing—talking to Claude, working 12 to 13 hours a day just to hit enter. The author says the real problem is not the AI-written code, but that nobody knows the system architecture and why these choices were made. Data engineering is singled out as a counterexample: that group was forced to understand the business from day one, and AI just saved them friction. I agree with this distinction, but I care more about the second part—among the friction saved, some was originally the understanding itself. You can let AI write, but architecture decisions and intent must be explainable by someone; otherwise, the next change will be a blind change.

Sources:


The tasks you have where the model fills things in yourself, do you dare add a phrase 'say I don't know if you don't know' and run it again? Dare, or check the effects first. See you tomorrow at 8 AM.

This issue selected 5 items from 72 pieces of information in the past 24 hours on X / Hacker News / GitHub Trending (collected hourly throughout the day, fact-checked, and compiled in the morning). Content is generated with LLM assistance, each item includes original source links, and important decisions should be cross-verified.

Like this brief? Get it by email

Daily AI coding picks at 8:00, plus a hands-on field-notes issue every Saturday. Written in Chinese.

This page is auto-generated by LLM aggregation; please cross-check with original sources.

Dev Breakfast · Adding 'Do not guess' cuts hallucination rate from 71% to 20% | Magic Tools | Magic Tools