Serious coding with LLMs. Lab notes 2026-10-01: Codex and Claude generate a shared analysis
The third ‘lab note’. More on the power and Achilles’ heel of “brute force trial and error with token pattern statistics”
Aggregated enterprise architecture wisdom
The third ‘lab note’. More on the power and Achilles’ heel of “brute force trial and error with token pattern statistics”
The second ‘lab note’. More on the power and Achilles’ heel of “brute force trial and error with token pattern statistics”
Before publishing about my very serious LLM-coding experiment, I am publishing a few shorter ‘lab notes’. This is the first one.
Coding with LLMs (Claude Code, OpenAI Codex) is often presented as the ‘killer app’ for Generative AI. But looking at data, it seems the one piece of the puzzle missing is actual cost. A quest into getting a less muddy picture about what is going on, w…
It turns out that AI has created a whole new language. Humans do not speak it, and they may even mistake it for talk about sex. But luckily Generative AI is able to translate it to something humans can understand (and where the sex doesn’t show up).
GPT-3o has done very well on the ARC-AGI-PUB benchmark. Sam Altman has also claimed OpenAI is confident that it can build Artificial General Intelligent (AGI). But that may be based on confusions around ‘learning’. On the difference between narrow, ge…
Google has announced ‘Willow’, a quantum computer that can calculate so fast it would take a supercomputer 10 septillion (a 10 with 25 zeros) years to do the same. But while the science is real and cool, the message is misleading. An explainer for non-…
One of the use cases I thought was reasonable to expect from ChatGPT and Friends (LLMs) was summarising. It turns out I was wrong. What ChatGPT isn’t summarising at all, it only looks like it. What it does is something else and that something else only…
Microsoft researchers published a very informative paper on their pretty smart way to let GenAI do ‘bad’ things (i.e. ‘jailbreaking’). They actually set two aspects of the fundamental operation of these models against each other.
If we ask GPT to get us “that poem that compares the loved one to a summer’s day” we want it to produce the actual Shakespeare Sonnet 18, not some confabulation. And it does. It has memorised this part of the training data. This is both sought-after an…
I came across a 2 minute video where Ilya Sutskever — OpenAI’s chief scientist — explains why he thinks current ‘token-prediction’ large language models will be able to become superhuman intelligences. How? Just ask them to act like one.
Google’s Gemini has arrived. Google has produced videos, a blog, a technical background paper, and more. According to Google: “Gemini surpasses state-of-the-art performance on a range of benchmarks including text and coding.”
But hidden in the grand …