Social Media Engagement: summer 2026

<div class = 'img-link'><a href = 'https://martinfowler.com/articles/2026-social-traffic.html'><img src="https://imgproxy.citrusreader.com/NRDMVfeBQfEcXAl7eFdBnXywU7tYQlI_93LGBjjySIo/raw:1/aHR0cHM6Ly9tYXJ0aW5mb3dsZXIuY29tL2FydGljbGVzLzIwMjYtc29jaWFsLXRyYWZmaWMvY2FyZC5wbmc" width="350px"></a></div> <p>A quick survey of recent engagement of my posts on social media, indicating which service has by far the most engagement, and which service has seen a precipitous decline since early 2025.</p> <p><a class = 'more' href = 'https://martinfowler.com/articles/2026-social-traffic.html'>more…</a></p>

2026/9/9
阅读更多

Fragments: September 8

<p>Christian Catalini says we’re in a situation where we are <a href="https://catalini.com/ideas/economics-of-ai/">vastly reducing the cost of generating things, but not the cost of verifying them:</a>.</p> <blockquote> <p>This explains why the first major AI products appeared in chat, image generation, and code assistance. Not because these were the hardest human problems, but because their outputs were relatively easy to inspect. A user can judge the tone of a message, look at an image, or run a test on a piece of code. […] The old automation boundary was routine versus non-routine work. The new boundary is increasingly measurable versus non-measurable work.</p> </blockquote> <p>The issue is then over how well you can measure something. In our profession, we know there’s a big difference between how many lines of code we write and how productive we are, and we’ve seen a regular <a href="https://martinfowler.com/bliki/CannotMeasureProductivity.html">failure to understand how to measure productivity</a>. Too much of what makes work effective is subject to either slow feedback loops or assessments that require subtle judgment. The danger is that people use lots AI automation while using incomplete measurements of its effectiveness, leading to short-term dashboards going up, but disaster in longer time-scales. He refers to these illusory short-term gains as <strong>counterfeit utility</strong>.</p> <blockquote> <p>Scale this across companies and institutions and the result is a <strong>Hollow Economy</strong>: extraordinary measured activity sitting on top of weakening human capability, hidden technical debt, correlated errors, and outcomes that nobody can confidently stand behind.</p> </blockquote> <p>Another highlight in the article was his advice to “build a history of decisions, not a gallery of outputs”. The point is that with AI we can all build really impressive things, but our value lies in the judgment that we’ve formed. It reminds me of how math problems were marked at school. We weren’t just marked on getting the final answer, we were also marked based on our reasoning process.</p> <p>He uses the OpenAI–Hugging Face incident as an illustration of this gap between generation and verification. He criticizes those who anthropomorphize the agents involved in the attack. By doing so we focus on the behavior of the AI agents, but instead we should focus on the financial incentives that created them and the environment they are operating in.</p> <blockquote> <p>Labs are locked in a race. The training run is where the money goes, and RL optimizes exactly what you score. The runs were scored on capability. They were not scored on “did not poison the Artifactory cache.”</p> </blockquote> <p>I assert that the organizations that build and run agents are responsible for <em>everything</em> those agents do, whether that behavior is intended or emergent. If they reap counterfeit utility by neglecting verification, they must face consequences: legal, financial, and if necessary: criminal. To deal effectively with AI, we need to change the incentives involved to ensure people invest more in verification than they do in generation. Otherwise we are driving a car that has a powerful engine, but weak brakes.</p> <p> ❄                ❄                ❄                ❄                ❄</p> <p>Brian Cantrill relates how <a href="https://bcantrill.dtrace.org/2026/09/05/the-revolt-of-the-reader/">readers are exasperated with “writers” using LLMs</a>.</p> <blockquote> <p>To those who read broadly, the hand of the LLM is so clear it’s as if the writer’s intellectual fly is open. In fact, it’s so jarring that I have to believe that those writing with LLMs are either not reading enough to see the LLM’s obvious structural tells — or (and?) they aren’t even reading their own content. (A confession: with particularly egregious pieces, I have fantasized about sentencing the author to read them aloud, certain that they themselves will be unable to endure the slop that they are foisting upon the rest of us.)</p> </blockquote> <p>He points out that readers do care about this, a survey found 78% of readers stop immediately once they sense something is the work a stochastic parrot, and 71% go on to blacklist the writer. It’s not the polish, it’s the authenticity that counts. Readers will always prefer the clumsy voice of the author over the gloss of an LLM’s whispering.</p> <p>Cantrill reports good success with using <a href="https://www.pangram.com/blog/pangram-4-technical">Pangram</a> to detect AI writing. I confess I’m a bit wary, do I really trust anyone’s judgment to disentangle LLM-voice from changes in generation and context? Maybe people steeped in Silicon Valley culture authentically speak in LLM-voice these days. Sadly for them, to be misclassified by their readers as an LLM is just as bad as using the damn things.</p> <p> ❄                ❄                ❄                ❄                ❄</p> <p>One of the dirty non-secrets about LLMs is that they were trained on a vast corpus of writing, without consulting the authors of that writing to see if they were cool with it. Individual authors like me can’t do a great deal about it, so are easy to ignore, but music companies aren’t exactly known for taking this kind of thing lying down. So they are <a href="https://www.theguardian.com/business/2026/aug/31/aanthropic-sued-alleged-theft-songs-ai-train-claude">suing over the use song lyrics for LLM training</a>.</p> <blockquote> <p>Sony Music Publishing and Warner Chappell, music publishers who manage the copyright of songs on behalf of songwriters and composers, are seeking damages for alleged misuse of “tens of thousands” of copyrighted works by Anthropic. […] The plaintiffs claim they are victims of “one of the largest and most blatant ongoing thefts of intellectual property in history”.</p> </blockquote> <p>Looking at it a broader societal point of view, there is an argument that the benefits of LLMs could be worth far more than any losses to us authors. But we should not forget that these tools are built on a foundation they used without our consent, and that should be taken into account as we regulate these tools and the fruits they provide.</p> <p> ❄                ❄                ❄                ❄                ❄</p> <p><a href="https://x.com/Steve_Yegge/status/2094947586373505122">Steve Yegge:</a></p> <blockquote> <p>All models, no matter how smart, will eventually build systems that they can no longer understand or maintain, if you let them. Fable 5 finally outbuilt itself, and flailed on me for a week. Fable 5.1 looks like it will fix it. For now. But you have to keep an iron grip on system size, or it’ll run away from you.</p> </blockquote> <p> ❄                ❄                ❄                ❄                ❄</p> <p>I was going through some slightly-related work and discovered that the Creating Passionate Users blog had disappeared from the internet (and has been gone since maybe a year ago). For those who don’t know, <strong>Creating Passionate Users</strong> was one of the treasures of the Golden Age of internet blogging. It was the work of <a href="https://en.wikipedia.org/wiki/Kathy_Sierra">Kathy Sierra</a>, also known for co-creating the “Head First” series of computer books. It talked about user experience, and remains some of the best writing on the topic, full of sparkling insights that greatly influenced my thinking, as well as many folks more engaged on user-experience work.</p> <p>Sadly not just was the blog ahead of time in its content, it was a harbinger of the darker side of the internet, as Kathy came under attack from a particularly virulent form of <a href="https://martinfowler.com/bliki/NetNastiness.html">Net Nastiness</a>. That led her to retreat from active participation on the web, and we’ve missed her ever since.</p> <p>Fortunately the Wayback Machine did its great duty, and <a href="https://web.archive.org/web/20250909191913/https://headrush.typepad.com/creating_passionate_users/">we can still read its snapshot</a>. I’ve often thought that, if I had a clone to spare, I’d like to create a guided tour of Creating Passionate Users to help readers today read that excellent material. (And if you’re reading this Kathy, and want it still hosted on the web, I’d be delighted to.)</p> <p> ❄                ❄                ❄                ❄                ❄</p> <p><a href="https://bsky.app/profile/simonwillison.net/post/3muwzopsoqs26">Simon Willison:</a></p> <blockquote> <p>“I don’t know the answer myself, but I asked a blowhard I know and he took a wild guess, here’s what he said: “</p> <p>How I interpret pasted replies from an LLM in online conversations</p> </blockquote> <p> ❄                ❄                ❄                ❄                ❄</p> <p>Jessica Kerr loves the feeling of being part of a team of people that learns from each other and from the codebase they are building as extensions of themselves - she incorporates the term <a href="https://jessitron.com/2018/04/15/the-origins-of-opera-and-the-future-of-programming/">symmathesy</a> for this: a learning system composed of learning parts (both the people and the code).</p> <p><a href="https://jessitron.com/2026/08/30/who-are-we-now/">“But now agents!”</a></p> <blockquote> <p>There was a turning point last year where I noticed that not only are they useful, it is irresponsible not to use them, at least in conjunction with my own code. They’re more thorough, as well as faster. How am I supposed to be responsible for this system, when I don’t understand each line of code?</p> </blockquote> <p>She has a habit of digging out old terms and ideas and applying them to our digital world. To frame what’s happening, she digs out two bits of latin</p> <blockquote> <ul> <li>Verum Factum: I made it, so I get it</li> <li>Vexationes Artium: Put it to the test [i.e. experiments]</li> </ul> </blockquote> <p>Agents can’t have Verum Factum knowledge, since it’s gone once their context window clears. They can use Vexationes Artium, running tests to see if something is working.</p> <blockquote> <p>If we want agents to write working, reliable code for us, we have to double down, 10x down on our objective verification. We need to vexate that code in artful ways. And we have the agent help us with that, with its thoroughness.</p> </blockquote> <p>This is, of course, true of those building these AI models - they certainly don’t have a Verum Factum knowledge of how they work, all they can do is come up with artful vexations to figure out what might be going on in there.</p> <p>What does that mean for us humans? Kerr says The Enlightenment elevated the idea that reason was the special quality of mankind. But now we’ve built machines that can reason. We need to focus instead on human qualities that the machines don’t have. Imagination is more important to us now than reason. And the essence of our humanity is in our relationships with other people.</p> <p>This material was put together for a conference talk, it’s available in <a href="https://www.youtube.com/watch?v=Nvr0klVxw5c">video</a> and in <a href="https://jessitron.com/2026/08/30/who-are-we-now/">text</a> form. Generally I far prefer text, but there’s so many delights in her talk that I found it worth watching. (Such as her current definition of legacy software, as code that wasn’t built with agents in mind - and is thus lacking in testability and consistency.)</p> <p> ❄                ❄                ❄                ❄                ❄</p> <p>Jim Gumbly outlines the decision space for <a href="https://www.jimgumbley.com/blog/timing-tradeoff-skynet-timeline.html">three people who are making vital judgments</a> about controls and regulations for AI: Sam Altman, Dario Amodei, and David Sacks (for US government).</p> <blockquote> <p>Sacks worries that binding oversight could weaken American leadership and the financing that sustains it. Sam and Dario face commercial incentives to keep producing more capable models. The value of those models increasingly comes from persistence, cooperation and permission to act. Those same properties increase the demands on containment and oversight. Some advances also make the evidence harder to interpret.</p> <p>Each decision can be understandable on its own. Their combined effect can still move deployment ahead of our ability to control it.</p> </blockquote> <p>A complication for all this is the issue of observability of these models.</p> <blockquote> <p>OpenAI reports that Astra is better aligned on its evaluations while also being less monitorable. Its traces can be shorter and less informative. In adversarial tests, where it was instructed to evade oversight, it could conceal some behaviour from monitors.</p> <p>[…]</p> <p>That matters to the familiar software cycle of releasing something, observing what happens and improving it. The cycle depends on the quality of the observations. Fewer warning flags are reassuring only to the extent that the warning system remains capable of detecting the relevant failures.</p> </blockquote> <p> ❄                ❄                ❄                ❄                ❄</p> <p>There’s an El Niño year coming up, and The Grauniad reports that climate scientists predict <a href="https://www.theguardian.com/environment/2026/sep/03/el-nino-supersizing-before-our-eyes-united-nations?CMP=share_btn_url">this El Niño is going to be a spectacularly hot one.</a></p> <blockquote> <p>The most recent data, from Monday, shows the temperature of the ocean at the heart of El Niño at 2.6C above the 30-year average. That is already close to the highest anomaly ever recorded in the satellite data era, 3.1C in 2015, with months to go before the peak is expected.</p> <p>That peak is forecast to reach about 4C in November, according to the average of 14 different models. Data from analysis of corals, tree rings and historical documents suggest no El Niño has reached this level in the last millennium, said Zeke Hausfather, a climate analyst.</p> </blockquote> <p>If these forecasts end up being accurate, will this make a difference to how seriously people are taking the climate crisis?</p> <!-- LocalWords: symmathesy -->

2026/9/8
阅读更多

Do you even need a presentation?

<div class = 'img-link'><a href = 'https://martinfowler.com/articles/never-send-slides/need-presentation.html'><img src="https://imgproxy.citrusreader.com/8iVx0eRtn5wuJxipycPI24vDXnpz4ODe0OjzNWXd2fU/raw:1/aHR0cHM6Ly9tYXJ0aW5mb3dsZXIuY29tL2FydGljbGVzL25ldmVyLXNlbmQtc2xpZGVzL2luZGV4L2NhcmQucG5n" width></a></div> <p>Like me, <b class = 'author'>Sumeet Gayathri Moghe</b> is tired of poor presentations with bad slide decks. He's started to write a series of posts on how to avoid these calamities, beginning with a post that <a href = 'https://martinfowler.com/articles/never-send-slides/need-presentation.html'>questions whether a presentation is needed</a> at all. </p> <p><a class = 'more' href = 'https://martinfowler.com/articles/never-send-slides/need-presentation.html'>more…</a></p>

2026/9/8
阅读更多

Bliki: Paracelsus Maxim

<p><b>The difference between a medicine and a poison is dosage.</b></p> <p>Often we talk about certain habits, in programming or life, are good or bad. But few things are simple binaries. Some vary with context: reading a book is a good thing sitting in my garden, but not while driving my car. But another variable is dosage: a little pain-killer salves my headache, but too much will kill me.</p> <p>The importance of dosage was noticed by a 16th century Swiss physician called Paracelsus. His quote was originally in German “Alle Dinge sind Gift, und nichts ist ohne Gift; allein die Dosis macht, dass ein Ding kein Gift ist.” which (<a href="https://en.wikipedia.org/wiki/The_dose_makes_the_poison">according to Wikipedia</a>) translates as “All things are poison, and nothing is without poison; the dosage alone makes it so a thing is not a poison.” It's also known as “The dose makes the poison” or if you prefer your sayings in Latin “dosis sola facit venenum”.</p> <p>In programming, global data is a good example of the Paracelsus Maxim (as I like to call it). A little global data, especially when immutable, can be a handy way of propagating information that may needed anywhere in a program, but it quickly becomes dangerous if there is a lot of it about.</p> <p>This kind of thing crops up in lots of places. So when thinking about when things are good or bad, we should always ask “in what contexts” and “in what doses”?</p>

2026/9/2
阅读更多

An Accidental Blackboard

<div class = 'img-link'><a href = 'https://martinfowler.com/articles/exploring-gen-ai/an-accidental-blackboard.html'><img src="https://imgproxy.citrusreader.com/ZjYAdczJpG6_qePpaBpKbr14DtuDPehLSN3Ra7fXoj0/raw:1/aHR0cHM6Ly9tYXJ0aW5mb3dsZXIuY29tL2FydGljbGVzL2V4cGxvcmluZy1nZW4tYWkvZG9ua2V5LWNhcmQucG5n" width></a></div> <p><b class = 'author'>Giles Edwards-Alexander</b> reports that during an experiment to see how productive a team could be using fully agentic engineering practices, the team accidentally prompted the agents into creating a blackboard coordination system inside the git repository.</p> <p><a class = 'more' href = 'https://martinfowler.com/articles/exploring-gen-ai/an-accidental-blackboard.html'>more…</a></p>

2026/9/2
阅读更多

Maybe We Shouldn't Be Reviewing All This Code

<p class = 'precis'><b>TL;DR</b><br><i>Or, perhaps the problem isn't that AI has broken code review, maybe it’s that we've been using code review to solve the wrong problems</i></p> <img src="https://imgproxy.citrusreader.com/EUZuBENrFZGecXiREhjo0gDD2sEfx_132fVQJUyg4Pg/raw:1/aHR0cHM6Ly9tYXJ0aW5mb3dsZXIuY29tL3JhY2hlbHMtcmFtYmxpbmdzL2NhcmQucG5n"><p>I was on a panel recently with Brian Houck from DX at Code Remix, hosted by Moderne. It was one of the more interesting panels I’ve done, largely because we disagreed. As my colleague Martin Fowler says, panels are much more interesting when people disagree and both sides have a good argument. Brian and I definitely did.</p> <p>Brian has since written a thoughtful piece called <a href="https://newsletter.getdx.com/p/what-are-code-reviews-even-for"><em>What are code reviews even for?</em></a> He is clearly passionate about his position, and I am passionate enough about mine that I’m writing this response. To be clear, I think we mostly want the same things. I just don’t think code review is the best way to get them. Brian is lovely, by the way, and encouraged me to write this. But I’d be lying if I said I didn’t want you to think I’m right by the end :)</p> <p>So what were we disagreeing about?</p> <p>AI is producing more code than humans can realistically review. Brian cites some pretty striking numbers: at Meta, significant lines of code per human-landed diff reportedly increased 106% in a year, while DX’s own data shows median pull request size increasing 64%.</p> <p>His concern, which I share, is that simply automating code review away risks losing all the other things we use it for. Code review isn’t just about finding bugs. It’s how teams share knowledge, teach junior engineers, build collective ownership and spread architectural understanding.</p> <p>My question is: <strong>why are we waiting until code review to do all of those things?</strong></p> <p>I’ve never particularly liked pull requests as the centre of the software development process. Not because engineers shouldn’t look at each other’s code, but because I’ve always struggled with the idea that we should build something, finish it, package it up, throw it over to somebody else and <em>then</em> have the important conversation about whether we built the right thing in the right way.</p> <p>And don’t even get me started on merge conflicts. I’ve lost too many hours of my life.</p> <h2 id="shift-the-judgment-left"><strong>Shift the judgment left</strong></h2> <p>One of the principles I learned very early at Thoughtworks was to shorten feedback loops. If feedback is valuable, don’t remove it. Move it closer to the decision it is informing.</p> <p>Take the things we say code review gives us.</p> <p>If we want to <strong>explore alternative solutions</strong>, I’d rather do that before implementing one of them.</p> <p>If we want <strong>knowledge transfer</strong>, pair. Sitting next to someone, physically or virtually, while they reason through a problem teaches you far more than reading their completed solution afterwards.</p> <p>If we want <strong>junior engineers to learn how experienced engineers think</strong>, let them work with experienced engineers while they’re thinking. Pairing comes to mind again here, but teams could also do design sessions collectively with a whiteboard before they write (or instruct the agent to write) anything.</p> <p>If we want <strong>collective ownership</strong>, organise teams so people actually build and operate software collectively rather than relying on a pull request to tell everyone what somebody else has already built. For this again use pairing, mob programming, or team design sessions around whiteboard.</p> <p>If we want <strong>architectural alignment</strong>, design together (I won’t repeat myself about pairing and team design sessions, oh wait…) and then encode the important constraints as fitness functions.</p> <p>And if we’re reviewing code for formatting, linting, known security problems or things that can be deterministically tested, automate them. We really shouldn’t still be arguing about whitespace in 2026.</p> <p>Pair programming, trunk-based development, automated testing, static analysis, fitness functions and security scanning all move feedback earlier. Increasingly, agents can participate in those loops too, challenging designs, testing assumptions and continuously verifying what is being built, but the real thinking is coming from experienced humans and if we want that experience to benefit the whole team then we have to act like one much earlier than code review.</p> <h2 id="review-by-exception"><strong>Review by exception</strong></h2> <p>None of this means nobody ever reviews code. There are absolutely changes where I want another experienced human looking. An example would be a fundamental architectural change. Assuming we did a design session as a wider team, we might want to review the code as a team or agree it was implemented right, or discuss if we want to change anything. Other examples could be something crossing a sensitive security boundary, a change with a huge blast radius, an unfamiliar part of a critical system or simply something where the team says, “I’m not confident about this.”</p> <p>Those are exactly the places where human judgment is valuable, but that’s very different from requiring a human to inspect every change because that’s the ceremony we’ve historically used to create confidence.</p> <p>And we know now it’s not viable to continue down this path, hence why code review keeps coming up as an issue or a blocker. If an agent can produce ten times the code but every line eventually queues up waiting for a senior engineer to inspect it, we haven’t created a ten-times engineering organisation, we’ve created a big backlog and a new bottleneck.</p> <p>And I don’t think the answer is an AI agent pretending to be the human reviewer so we can preserve exactly the same process at higher speed. That’s automating the ceremony rather than questioning why the ceremony exists.</p> <p>There is one thing I do worry about in Brian’s argument, though. He talks about teams accumulating cognitive and intent debt: software grows while the humans responsible for it understand less and less about why it works the way it does. I think that’s a very real problem. I just don’t think mandatory pull requests are a particularly strong defence against it.</p> <p>If agents are going to produce substantially more of the implementation, we need to be much more deliberate about maintaining human understanding through collaborative design, pairing, good boundaries, executable architecture, shared operational responsibility and probably some practices we haven’t invented yet.</p> <p><strong>We need engineers to understand systems, not diffs.</strong></p> <p>Perhaps that’s what AI is exposing. We’ve spent years loading an extraordinary number of responsibilities onto the humble code review: quality gate, security check, architecture review, mentoring mechanism, knowledge-sharing system, ownership model.</p> <p>It worked, sort of, while humans could only produce code so quickly. That constraint is disappearing. So perhaps the question isn’t how we get the code reviewed faster. Perhaps it’s why we’re waiting until code review to have all the important conversations in the first place.</p>

2026/9/2
阅读更多

Fragments: September 1

<p>Like many readers, I’m wary of AI generated prose. Simon Wilison has written an <a href="https://tools.simonwillison.net/llm-cliche-highlighter#https://martinfowler.com/articles/making-data-ready-for-agentic-ai.html">LLM cliché highlighter</a> - paste in some text, or a URL, and it will flag various patterns common to LLMs. It references a <a href="https://en.wikipedia.org/wiki/Wikipedia:Signs_of_AI_writing">wikipedia page of signs of AI writing</a>. That page points out that:</p> <blockquote> <p>Humans are notoriously bad at distinguishing human and LLM-generated text. While research on humans’ abilities to detect AI-generated text is still limited, a 2025 study has shown that human ability to distinguish LLM text from human is no better than random chance. Another 2025 study on German theses has shown that humans managed a “recognition rate of 57% for AI texts and 64% for human-generated texts”.[</p> </blockquote> <p>Not just do I find myself repelled by prose with an LLM-voice, I also wonder how accurate my reaction is. I’m old enough to see all sorts of new tic-phrases appear, and in the past would just chalk it up to youngsters or airport business books. (Not to mention Americanisms, which I’ll get used to momentarily.)</p> <p> ❄                ❄                ❄                ❄                ❄</p> <p>NVIDIA’s technical blog reports on an <a href="https://developer.nvidia.com/blog/nvidia-avo-reaches-100-on-arc-agi-3-demonstrating-a-frontier-level-general-purpose-architecture-for-long-horizon-autonomous-agents/">Architecture for Long-Horizon Autonomous Agents</a>. Their research group used a combination of Claude Opus 5 and a harness called AVO, and used it first to do GPU kernel optimization and then a broader reasoning benchmark (ARC-AGI-3). Both of these were long-term tasks, for the kernel optimization the agent ran for seven days.</p> <blockquote> <p>AVO is designed to preserve progress beyond a single model context. Two mechanisms are particularly important: persistent memory and supervision.</p> <p>Persistent memory carries forward prior implementations, evaluation results, compiler and profiler outputs, and accumulated reasoning, allowing the agent to resume from the current state rather than repeatedly reconstructing the search.</p> <p>The supervisor monitors the broader trajectory for stagnation or repeated unproductive cycles and can redirect the main agent toward alternative strategies when needed. During the seven-day attention-kernel run, the main agent remained responsible for deciding what to inspect, change, test, and evaluate, while the supervisor helped maintain forward progress when the search plateaued.</p> </blockquote> <p>The team was encouraged that AVO did well at two different kinds of long-horizon tasks, indicating that it’s a general-purpose tool.</p> <p> ❄                ❄                ❄                ❄                ❄</p> <p><a href="https://x.com/mickeynp/status/2092525399209058394">Mickey Petersen</a>:</p> <blockquote> <p>MCP is SOAP for Zoomers.</p> </blockquote> <p> ❄                ❄                ❄                ❄                ❄</p> <p>Paul Stack writes that <a href="https://stack72.dev/ai-broke-the-assumptions-behind-ci/">AI Broke the Assumptions Behind CI</a>. Here’s his description of CI with agents.</p> <blockquote> <p>An agent writes a change, opens a PR, and CI picks it up instantly. The compile fails, the agent pushes a fix, CI picks it up instantly again. A test fails, another fix, another instant run. Each iteration is fast, but the agent is still discovering that its change doesn’t work only after it crosses the PR boundary. The feedback loop is in the wrong place regardless of how fast CI runs.</p> </blockquote> <p>He points out that all of this breaks the pipeline, because “CI” keeps failing, and advocates doing verification before the agent pushes. This is where I get to be the grumpy old guy, and point out that was <a href="https://martinfowler.com/articles/continuousIntegration.html#BuildingAFeatureWithContinuousIntegration">always how Continuous Integration works</a>. When I’m done with a change, first I pull (to get everyone else’s change since I started), I build and test locally, and if all is well I push and let the CI server do its thing. The only reason the CI server should fail is if there’s some funky mismatch between my machine and the CI server. Tests that take a while to run aren’t part of this loop, instead they are run further down the <a href="https://martinfowler.com/bliki/DeploymentPipeline.html">deployment pipeline</a>, downstream of CI. Any failures there imply missing tests in CI.</p> <p>(I’m being a bit unfair dumping on this article here. After all I could have filled a full working day correcting misleading descriptions of Continuous Integration for most of the last twenty years. Maybe I’m just after an excuse to point readers to the <a href="https://martinfowler.com/delivery.html">extensive range of articles</a> hosted here about what’s needed to get code from laptop to production.)</p> <p>Stack is right that we should question how the deployment pipelines should work with agents in play. He’s also right that CI with humans relies on them being disciplined to run commit tests locally before pushing to the CI server - and that we can (and should) automate that when using agents. I also don’t know more about his setup than what he’s written in his post, so there’s likely complications he faces that I don’t understand. But when thinking about designing pipelines it’s important to understand the principles that underlie <a href="https://martinfowler.com/bliki/ContinuousDelivery.html">Continuous Delivery</a>, understand how the practices really work, and understand why they are in place. Above all, Continuous Integration is a practice, not just the CI server. Yes, CI does conflate two jobs: executing verification and coordinating merges. But that’s the point: verification is a necessary part of merging if we want to retain a healthy <a href="https://martinfowler.com/articles/branching-patterns.html#mainline">mainline</a>.</p> <p> ❄                ❄                ❄                ❄                ❄</p> <p>Recently Noah Smith posted an article about how he was worried about an <a href="https://www.noahpinion.blog/p/heres-how-were-all-going-to-die">AI-generated super-virus savaging humanity</a>. It’s a worry I’ve heard a few times, seen as a greater concern than AI turning us into labradors or paper-clips. Claus Wilke, who works in the field, <a href="https://blog.genesmindsmachines.com/p/im-sorry-youre-not-going-to-die-from">isn’t so concerned</a>.</p> <blockquote> <p>Computational design of biological systems is unfathomably difficult. Experts who have dedicated their life to this topic routinely hit their head against the wall when nothing they try seems to work. PhD students in 2026 using state-of-the-art AI software are spending months or years trying to design simple peptide binders that inhibit some enzyme or pull down some protein, and the majority of their designs fail, or don’t express, or are toxic. But in Smith’s fictitious world a disgruntled teenager with no special training in biology can just solve a problem thousands of times more complicated than designing a peptide binder. The distance between where we are today and where we would have to be for Smith’s story to have any realism is enormous.</p> </blockquote> <p> ❄                ❄                ❄                ❄                ❄</p> <p>It seems that a couple of remarkably talented academics are experts in a staggeringly wide range of fields. <a href="https://arxiv.org/pdf/2606.02184">Or maybe they are just ghosts.</a></p> <blockquote> <p>These names do not exist. Elena Vasquez and Marcus Chen have appeared as volcano experts, astronauts, thriller protagonists, podcast hosts, and academic co-authors across hundreds of independently produced AI-generated documents, never having lived. We show that large language models do not merely default to high-probability individual names when generating fictional experts: they produce correlated character ensembles: pairs and trios whose co-occurrence rates far exceed chance and are consistent across independent generations.</p> </blockquote>

2026/9/1
阅读更多

Making Your Data Ready for Agentic AI

<div class = 'img-link'><a href = 'https://martinfowler.com/articles/making-data-ready-for-agentic-ai.html'><img src="https://imgproxy.citrusreader.com/BRFglycR2MTueYH9c4p154PLf7nRkTGdgA4_JBTJy18/raw:1/aHR0cHM6Ly9tYXJ0aW5mb3dsZXIuY29tL2FydGljbGVzL21ha2luZy1kYXRhLXJlYWR5LWZvci1hZ2VudGljLWFpL2NhcmQucG5n" width="350px"></a></div> <p>Lots of organizations are excited about what AI can do to streamline their processes, save money, and juice margins. But AI's capabilities are founded on the data that AI accesses, and for many organizations that foundation is little more than sand. <b class = 'author'>Pramod Sadalage</b> and <b class = 'author'>Prem Chandrasekaran</b> write about how to build a reliable foundation of data that can be accurate and trusted.</p> <p><a class = 'more' href = 'https://martinfowler.com/articles/making-data-ready-for-agentic-ai.html'>more…</a></p>

2026/8/27
阅读更多

Fragments: August 24

<p>I was listening to <a href="https://www.nytimes.com/2026/08/18/opinion/ezra-klein-podcast-helen-toner.html">Ezra Klein’s interview with Helen Toner</a> about the recent OpenAI hack of Hugging Face and the subsequent discovery that there were swarms of agents inside OpenAI doing unsanctioned activities. One of the points Klein made was that at no point did any of these (thousands of?) agents ever try to check in with a human</p> <blockquote> <p>[Klein:] So these message boards — you have however many A.I. agents posting hundreds of thousands of messages. At no point do they say: Hey, researchers, programmers, parents at OpenAI, Anthropic — do you want us coordinating with each other on this message board we have created in the innards of your systems?</p> </blockquote> <blockquote> <p>[Toner:] Or even F.Y.I., we have a message board we’re coordinating on in the innards of your system.</p> </blockquote> <p>Listening to that, another thing occurred to me - <em>none of these agents thought to rat the others out</em>. No “hey, some of the agents in here are doing sketchy things”, no sign of an AI whistleblower.</p> <p> ❄                ❄                ❄                ❄                ❄</p> <p>Is the AI bubble so big that the frontier companies like OpenAI and Anthropic have no way of becoming a viable business? If that’s the case, Bruce Schneier and Nathan Sanders <a href="https://www.schneier.com/blog/archives/2026/08/if-the-markets-reject-openai-and-anthropic-the-us-should-nationalize-them.html">have a possible path:</a></p> <blockquote> <p>Evidence suggests the market itself could reassess that these companies offer nothing of financial value. In that case, perhaps we can return them both to their original purposes. If these AI companies should fail in the financial markets, the US should nationalize them and convert them into national labs operated under democratic control that preserve their benefit to the public interest.</p> </blockquote> <p>Such an idea may strike many people, used to the laissez-faire free enterprise world of Silicon Valley, as sacrilege, disaster, even socialism. But the United States made world-beating technological progress through such institutions in the recent past. AT&amp;T was a quasi-government entity that led the world in telecommunications and electronics after the second world war.</p> <blockquote> <p>The US has a long, successful history of these kinds of institutions, which have produced world-shaping innovations in spaceflight, telecommunications, nuclear power and more. Congress currently manages a $200bn R&amp;D portfolio, within which frontier AI development is, arguably, a glaring gap.</p> </blockquote> <p> ❄                ❄                ❄                ❄                ❄</p> <p>Here’s a message for those readers who live in Massachusetts, just to the north of me, specifically in congressional district MA-06. I don’t usually endorse political candidates, but I’ve made an exception for <a href="https://bethfordemocracy.com/">Beth Anders-Beck</a>, who is running for that house district. I’ve known Beth for many years and have a high opinion of her smarts, wisdom, and compassion. They would make an excellent member of congress.</p> <p> ❄                ❄                ❄                ❄                ❄</p> <p>Kevlin Henney posts “one weird trick” for <a href="https://kevlinhenney.medium.com/streamline-your-linkedin-experience-with-this-one-weird-trick-041daa16777a">deciding when to skip reading LinkedIn posts</a>, essentially by identifying a common pattern for skippable posts:</p> <ol> <li>Post is too long</li> <li>Contains a (crummy) info-graphic</li> <li>No voice of poster (instead “aspiring anodyne anonymity”</li> </ol> <p>It seems like a good approach. I, however, have a simpler one - skip all LinkedIn posts.</p> <p> ❄                ❄                ❄                ❄                ❄</p> <p>Bartosz Ocytko has detailed and thoughtful post about the usage of <a href="https://engineering.zalando.com/posts/2026/08/agentic-engineering-at-zalando-a-snapshot.html">agentic programming at Zalando</a>. Like most companies I hear from, they are convinced of the value of agentic programming but still exploring how best to do it. One notable step they’ve taken is building platforms to act as a clear portal for API access and tools to support chat UI and CLI. This allows them better support good security practices and to monitor usage of models.</p> <p>They have seen signs of agentic programming increasing the complexity of codebases, including leading to larger commit messages.</p> <p>The write-up spends a lot of time on knowledge sharing, how to pass on skills, and the support of experiments.</p> <blockquote> <p>With &gt;200 teams innovating and broadly exploring the ecosystem, the question arises whether and when to converge. We believe it’s way too early for this. While agentic engineering practices are still in their early stages, our key objective is transparency and exchange across teams.</p> </blockquote> <p>I was struck by their use of an LLM to assess the risk of pull-requests. Those with a low risk of rollout can be auto-approved, reducing lead time by 20-40%. An interesting consequence of this is that it encouraged folks to split pull-requests so low risk portions can take advantage of the fast approval. Any changes to configurations are automatically made high-risk, which they feel protects them from common outage traps.</p> <p>They repeat the common thread that the value of AI depends greatly on underlying skills.</p> <blockquote> <p>Like anyone in the industry we observe how AI amplifies the good and bad practices across our organization. Teams that get carried away with agentic engineering end up with large PRs that discourage reviewers and slow down delivery until a team adjusts their practices.</p> </blockquote> <p> ❄                ❄                ❄                ❄                ❄</p> <p>Julia Curlee was a senior intelligence official in the White House. She had served under administrations of both parties, been the briefer for Vice President Pence, and on the National Security Council under Biden. She writes an absorbing account of her relationship with Pence and shares observations about the changes to the intelligence community under the current administration, including <a href="https://www.theatlantic.com/magazine/2026/10/trump-white-house-transgender-mike-pence/688284/?gift=zGsHlQiVhVhk3cFVqv--gzIqOYZAxuFeY4CtPlqqfx4&amp;utm_source=copy-link&amp;utm_medium=social&amp;utm_campaign=share">recent events at the CIA</a> (gift link)</p> <blockquote> <p>The agency has been gutted as part of a deliberate plan, the director of the Office of Management and Budget once boasted, to put the people who defend our country “in trauma.” Analysts have been fired in public or questioned by the FBI; decade-old assessments have been denounced by the CIA director in the press. The president calls analysis “virtual treason” when it contradicts his preferred reality, and uses the CIA to undermine public confidence in American elections.</p> <p>Fear has done its work. Irreplaceable officers with crucial language and technical skills, and decades of experience, have walked out the door. Those who remain within an agency built to deliver hard truths are being muzzled.</p> </blockquote> <p>For a worthwhile sample of her analysis, read this evaluation of the <a href="https://www.lawfaremedia.org/article/fighting-while-talking--the-iran-war-enters-its-bargaining-phase">current bargaining between the US and Iran</a></p> <blockquote> <p>Most wars do not end in “unconditional surrender.” They end when both sides accept terms. Paul Pillar’s classic study of war termination, “Negotiating Peace,” treats combat and diplomacy as a single process: Each side fights to improve the terms it can demand at the table, and talks to lock in what the fighting has won.</p> </blockquote> <p>She continued to serve the second Trump administration even though they knew she was trans, until her position was made public.</p> <p>Autocrats seem appealing, with the promise to get things done without the ponderous constraints of rule of law or bureaucratic procedure. There are occasional “Good Emperors” who raise people based on merit, but more often such power attracts corruption, nepotism, and toadies.</p> <blockquote> <p>Flailing regimes dehumanize minorities to distract from their failures. When the economy collapses or a war goes badly, they find a tiny group of people, make them the enemy within, and rally the country against them. This is how it’s gone in Iran. Hungary. Russia. I wrote PDBs about it. This will not stop with trans people. It never has.</p> </blockquote>

2026/8/24
阅读更多

Citizens Build, Agents Execute, Experts Govern

<p class = 'precis'><b>TL;DR</b><br><i>Why building an app over the weekend isn't the same as building enterprise software</i></p> <img src="https://imgproxy.citrusreader.com/EUZuBENrFZGecXiREhjo0gDD2sEfx_132fVQJUyg4Pg/raw:1/aHR0cHM6Ly9tYXJ0aW5mb3dsZXIuY29tL3JhY2hlbHMtcmFtYmxpbmdzL2NhcmQucG5n"><p>I’ve noticed an interesting gap opening up over the last six months. It isn’t really a gap in technology. It’s a gap in what different people think software engineering actually is.</p> <p>The conversation usually starts the same way. A non-techie, maybe an executive, tells me about something they’ve built over the weekend. Sometimes it’s a chatbot. Sometimes it’s an internal workflow. Sometimes it’s a surprisingly polished application that solves a real business problem. They’re excited, and they should be. Twelve months ago they probably couldn’t have built it at all. Then comes the question.</p> <p>“If AI can do this now, why aren’t our engineering teams delivering ten times faster?”</p> <p>It’s a perfectly reasonable question, after all we’ve all seen the demos. The first thing that would come to my head is “you don’t know what it takes to build enterprise grade software”. But then I think about what I mean and how to explain it to a non-technical person without sounding super patronising. And then it hit me, we did this to ourselves. We’ve spent so many years banging on about how to write good software that everyone has assumed writing software is the same as software engineering.</p> <p>The application someone builds over the weekend is real software. It likely solves a real problem or demonstrates an idea. Sometimes it’s genuinely impressive. I don’t want to diminish that because I think one of the most exciting things AI has done is dramatically increase the number of people who can turn ideas into working software. That’s cool, I totally get it. The first apps and “hello worlds” I ever built excited me enough to choose this as an actual career so the excitement is real and I don’t want to temper it too much.</p> <p>But your first hello world, which these days can be an entire app with all kinds of features, is very, very (extra very on purpose) different from introducing software into a production environment in a highly regulated enterprise, as an example. But why?</p> <p>The moment that application becomes something the business depends on, the questions change completely. Is customer data protected? What happens when a dependency fails? Can someone else understand this system in two years’ time? Will it survive an audit? Can it cope with a thousand times more users than it has today, what about millions in one day? How will we know something is wrong before our customers do? Those questions don’t show up in a demo or in the build phase at all unless an experienced engineer is in the room. I certainly wasn’t asking them when I was building my first apps. I only cared about features!</p> <p>This is where experienced engineers become more important, not less. Not because they’re the only people who can build the software anymore, but because they have the judgement to know whether we can trust it: whether the design is good, the risks are understood, and the thing that works today won’t become somebody else’s nightmare six months from now.</p> <p><a href="https://martinfowler.com/bliki/FutureOfSoftwareDevelopment.html">At FOSE</a> a few weeks ago, we spent surprisingly little time talking about coding. We talked about whether code was still the source of truth, and occasionally about how much we missed writing it, but mostly we talked about design, architecture, governance, learning and judgement. One team described spending the day designing a specification, letting agents work overnight and reviewing the results the next morning. The interesting bit for me wasn’t the overnight pipeline, cool as that was. It was what the humans were doing: deciding what good looked like, making trade-offs and judging whether what came back was actually what they wanted. We also kept coming back to good design, because it turns out that when agents can generate lots of code very quickly, good design matters more, not less.</p> <p>That made me wonder whether we’ve been thinking about scarcity in the wrong way. We’ve spent decades optimising around people who can write code because they were scarce and expensive. I’m not convinced that was ever the real scarcity, but that’s probably another ramble. What feels scarce now is good engineering judgement: knowing what good looks like, understanding the risks and knowing when something that works is actually safe to trust in production. Because software doesn’t exist to be built. It exists to run in production and safely solve the problem it was created for. <strong>Organisations don’t run on code. They run on trust.</strong></p> <p>A few months ago I found myself saying something in a conversation almost without thinking.</p> <p><strong>Citizens build. Agents execute. Experts govern.</strong></p> <p>It sounded cool and I thought marketing would like it, so I wrote it down. Then I left it alone for a while. The funny thing about writing these ramblings is that I don’t know whether I believe something until I’ve let it bounce around in my head for a while and also said it to other people I trust like senior engineers at Thoughtworks. Sometimes I come back convinced I was talking nonsense. Occasionally I realise there was something more interesting hiding underneath. This was one of those occasions where the latter was true.</p> <p>At first I thought I was talking about roles. Citizens build software (essentially non-engineers). Agents write the code. Engineers become governors. But I don’t actually think that’s what I meant. I think I was talking about where value is moving. AI has given everyone a new way to express their ideas. The execution is increasingly handled by agents. They write the code, refactor it, generate tests, fix bugs and iterate at a speed that simply wasn’t possible before. But neither of those things reduces the need for expertise.</p> <p>In fact, I think it does exactly the opposite. When everyone can create software, somebody still has to decide whether that software deserves to exist inside an enterprise system in PRODUCTION. Somebody still has to think about architecture. Security. Resilience. Operability. Compliance. Cost. The boring stuff that nobody gets excited about in a demo but that becomes painfully important the first time a customer can’t log in or an auditor comes knocking.</p> <p>That’s why I don’t think experienced engineers become less important. I think they become dramatically more leveraged. Their job shifts from building every feature themselves to creating the environment in which thousands of features can be built safely by other people and by agents. They become the people who design the guardrails, the platforms, the engineering practices and the feedback loops that allow everyone else to move quickly without creating chaos.</p> <p>Perhaps that’s the future software organisation. Not one where everyone becomes a software engineer. Not one where software engineers disappear. One where almost anyone can create software, agents increasingly execute it, and engineering expertise becomes the thing that allows all of that creativity to scale safely. And to be clear I do not mean people build stuff and throw it to engineers to fix, that is a total antipattern for another ramble.</p> <p>Perhaps that’s why the executives and engineers I’ve been speaking to sometimes sound as though they’re describing completely different futures. The executive sees that anyone can now build software. The engineer sees that somebody still has to live with it. Both are right. They’re simply looking at different parts of the same system we have to solve to create whatever the future actually ends up being.</p>

2026/8/19
阅读更多

推荐订阅