When you use a tool,
... you usually know what it can do. A knife cuts, a hammer drives in nails, a pen draws a line.
Of course, tools can be specialized for particular tasks (e.g., a fish knife or a ball peen hammer), or can be floridly multi-purpose (like your favorite computer or a Swiss Army Knife like that above).
This raises a great question about what's your mental model of an AI system?
Classically, AI systems were fairly specific in what they could do. In the great AI expert systems of the past, they could do exactly one thing and do it very, very well. But with the rise of modern AI, especially LLMs, things have changed, and now AI systems can do a wide variety of tasks.
For us, the question is: what are those tasks?
This is the problem of knowing what's possible with an AI.
People have written about the "jagged edge of AI" as being a difference in the capability of a single AI system. (For details, see "Navigating the jagged technological frontier.") Your AI can do some complicated things really well but then it will fail on similar tasks that are fairly straightforward, giving rise to the sense that the edge of competence is very jagged.
That's unfortunate because it's tough to tell what your AI will be able to do.
But there's also a jagged edge between different AI systems. That is, just because one AI system can do a particular thing, it's unclear that another, different AI can do the same thing.
Here's an example... figuring out how old a tree is.
While out on a hike this weekend, I came across an exposed tree stump that had been cut down. It had shown a bunch of tree rings, and I was wondering how old the tree was. This, I thought, was the perfect application for an AI.
So I gave three popular AI systems the task of counting the number of tree rings on a stump. As you can see the results between the three AI systems are very, very different. I prompted each by uploading the image above and then giving it the task of [how old is this tree] letting it figure out the method by which it should count tree rings.
Claude (Opus 5, medium) did a reasonably good job that looks very impressive but it's not actually very accurate even though it gives the impression of high accuracy. The response has lots of technical language and botanical terms. The image analysis looks very precise, grainy, and scientific. The whole this is impressive... but wrong.
ChatGPT 5.6 did a pretty good job, although if you look carefully you'll see that the numbers of the tree rings are not quite right. Things don't line up the way they should.
Gemini did a pretty terrible job. If you look at the numbers on the tree rings, you can see that the zero is actually out about five years after the tree started to grow. If you look carefully there are rings that are labeled both 5 and 10 or if you look a little bit further out, there are rings labeled both 30 and 25.
What's up with that?
(Interestingly, the first time I asked Gemini for the rings on the image, it said it created the image, but didn't REALLY. But by asking for it to re-do the image, it coughed up this image, but it also changed its initial estimate of 52 years old to 68. I don't think it's being sycophantic, just nondeterministic!)
[generate an exploded parts diagram of a modern racing bicycle]
![]() |
| Gemini's version of a bicycle, not bad, but also NOT exploded. |
ChatGPT:
![]() |
| ChatGPT's bike is more exploded that Gemini's. |


























