What’s missing

Arles, August, 2026
I’m pedantic about apostrophes, so of course this struck me. And since it generally indicates that a word is contracted, I wondered how the owners of the café came up with the name. But then I looked it up, and discovered that it can also mean “a figure of speech by which the writer suddenly breaks off from the previous method of his discourse, and addresses, in the second person, some person or thing, absent of present”.
Example: “Milton’s apostrophe to Light at the beginning of the Third Book of Paradise Lost”.“
Quote of the Day
”I’m not a heavy drinker, I can sometimes go for hours without touching a drop.”
- Noel Coward
Musical alternative to the morning’s radio news
Alison Krauss | Baby, Now That I’ve Found You
Long Read of the Day
Slop-vestigation and the digital pantograph
Really interesting essay by Robin Sloan triggered by reading the investigation of the Hugging Face hack.
The OpenAI-Hugging Face incident remains THE fascinating event of the summer, maybe the year; deeper investigation has revealed its rich, strange structure.
But, notice:
Over the course of this investigation, OpenAI provided us with the dump of ~1.2 million entries from the main message board and the dataset of ~1300 transcripts we describe below, as well as free API credits for GPT-5.6 Sol for analysis.
How do you make sense of 1.2 million agent messages and 1300 very long LLM agent activity transcripts? With another LLM, of course.
This is a pattern that recurs in this domain. Assembling training data at the scale required by 2020s-era models, no researcher can “read it all”. So, you either (1) don’t bother, or (2) use another LLM to review and filter the data. You can, in principle, use other kinds of models — simpler classifiers — but, increasingly, the kinds of judgments you need to make require the richness of an LLM.
Anthropic’s Insights tool, likewise, uses Claude to read and categorize millions (billions?) of transcripts of people’s interactions with Claude. In addition to making this huge heap of data legible at all, the “LLM in the middle” acts as a privacy buffer: researchers read only Claude-generated summaries, not the original interactions.
I’ve come to think of this as “using tongs”, in the sense of a tool that allows you to manipulate material that you otherwise couldn’t…
Brilliant essay, which illuminates the Catch-22-type dilemma in which we’re enmeshed.
Linkblog
Something I noticed, while drinking from the Internet firehose.

If you were an early user of the Internet in the 1980s, then UseNet was the place to be. Which is why this searchable archive is an interesting re-apparition.
This Blog is also available as an email three days a week. If you think that might suit you better, why not subscribe? One email on Mondays, Wednesdays and Fridays delivered to your inbox at 5am UK time. It’s free, and you can always unsubscribe if you conclude your inbox is full enough already!