Perspective

Draw a Banana: How AI Models Work

How these systems predict, what they see, and where bias comes from

Kristina Agustin
August 30, 2026
~7 min read
Back to Chart Room
Share
Author: Kristina AgustinPublished by: Southern Sky AI

A friend of mine went to hear Tea Uglow speak recently. Uglow was a founding member of Google's Creative Lab, built its teams in London and then Sydney across seventeen years, left in 2023, and has been working with AI since 2016.

She opened by showing the audience an old machine learning game. A prompt appears: draw a banana. Draw a cat. You draw one, and the system decides whether you have come close enough to what it expects a banana to look like. If you have, you go through to the next level. If you have not, you try again. My friend's account of Uglow's line was that she thought hers was pretty good, the machine disagreed, and you keep drawing until yours matches theirs.

That, Uglow said, is what a language model does. It is a predictive system. It carries a picture of what is normal and likely, and it moves you toward it.

Then she widened it. Anyone whose circumstances sit outside whatever the system has learned to treat as ordinary is served less well by it, and Uglow has spent a career on who that leaves out. Then she made it universal. Even if you fit the expected picture comfortably today, she said, you are going to age. Everyone ends up outside it eventually.

Set that against the pace of the last two years. There is a new model most weeks, a new capability, a tool that supersedes something you had only started using. Keeping up with the releases is exhausting, and it is mostly not where the advantage sits. I made that case in the first episode of the Learn series, and if anything it has accelerated since.

The principles of how these systems work stay steady. Once you have those properly, you can evaluate whatever arrives next, because you're reading the technology underneath rather than the announcements around it.

The most valuable part of every AI course I've done has been the foundational material, and that's where the light-bulb moments came from for me. Tools are tools. Each one is a window into a model, so they're rarely the competing technologies the marketing makes them out to be, and mostly they're different ways of working with the same intelligence.

Most of the AI training on offer is about prompts and tools. Very little of it goes to what's happening when you press enter.

01

One. The board

When I started my AI consultant certification with Innovating with AI in October 2024, Rob Howard explained these systems using a game from The Price Is Right.

A contestant drops a disc into the top of a Plinko board. It bounces down through rows of pegs and finishes in one of the slots along the bottom. Any single drop is unpredictable. Drop enough discs and a pattern appears.

Individually each drop looks like chance. Collectively the discs settle into the same shape, and they do not spread evenly across the slots. They collect in a few near the middle, and the outer slots take one rarely.

02

Two. Language

A language model breaks text into tokens, which are chunks of language, sometimes a whole word and often part of one, and predicts the next token from probability, context and the patterns it saw in training.

That is the board, in words. Each next token is a disc falling, and the shape of the board is everything the model read. Some continuations are far more probable than others because they turned up far more often, so what comes back to you is the answer sitting where most of the discs collect.

A model can give you an answer that is fluent, confident, well organised and entirely ordinary. It is doing what it was built to do. It is producing the most likely response, and most likely is a separate quality from correct, and separate again from the answer your particular situation needs.

The longer version is Episode 2, five minutes on what these systems are, from Alan Turing's original question through to today.

03

Three. Vision

The same logic runs through images, and it is easier to see there because pixels are more concrete than words.

One of the classic tests is the old question: is it a cat or is it a dog? To a machine, an image is a grid of pixels made of colour and light values. A computer vision model learns to recognise patterns in those pixels, in shapes and edges and textures, and from those patterns works out what it is looking at. It is trained on millions of labelled images, learning what makes a cat a cat and a dog a dog, even when the photographs vary in lighting, angle and background.

Over time it does more than spot animals. It can identify defects on a hull, monitor whether safety gear is being worn, or scan containers for pests. Ports and fleets are already running this at scale, and there are worked examples across our industry in Episode 3, from autonomous shipping to the automated terminal in Singapore.

The logic is the same as with language. Given enough examples, the system learns what to expect and where. Same board, same discs, different medium.

04

Four. Bias, which is the same mechanism

Put those three together and bias stops being a mystery, or an accusation, or something careless engineers left in.

A language model learns what is probable from what it read. A vision model learns what a cat looks like from the cats it was shown. Both are assembling a picture of the ordinary out of a body of examples, and both are at their weakest when a case is uncommon, because uncommon cases are exactly what those examples contained least of.

As I put it in Episode 4, these models are trained on the entirety of human history, so by nature they are shaped by our past, and our past has not always been fair. History carries deep-rooted biases across gender, race, class and geography, and a system built to reproduce what was common will reproduce those alongside everything else. Not because anyone intended it. Because that is what shaped the board.

So bias is a design property rather than a defect. It comes from the same mechanism that makes these systems useful, and whatever lets a model finish your sentence sensibly is the same thing that makes it reach for the ordinary.

Considerable work goes into reducing it, in how training data is curated, in fine-tuning, in the guardrails around what a model will say, and that work does help. What it does not do is remove the property, because the property is where the capability comes from. So bias gets managed rather than solved, which is a different job and a standing one.

Bias usually gets taught as a compliance module, a box sitting next to privacy and security, something you acknowledge on a slide and move past. I'd put it much earlier than that, inside the technical description of how these systems work, because that's where it comes from. If you treat it as an ethics item you'll write a policy about a fault nobody put there, and the thing itself carries on underneath it.

A system built around the ordinary serves the unusual case less well, and every one of us is an unusual case at some point, including the person who fits neatly today and will not in twenty years. That was Uglow's point.

Closer to home, whatever you know that few others know is, by definition, uncommon in what these systems read. The specialised knowledge that makes your work valuable sits out in the slots where hardly any discs fall. These tools are at their weakest exactly where your expertise is deepest, and that weakness comes with no obvious tell, because a general answer to a specialist question arrives with the same fluency as a correct one.

So audit what comes back. Watch for the answer that is smooth and average. Keep somebody with real knowledge of the subject between the output and anything that matters, because they are the only one who can see the difference. If you are setting that up from scratch, Episode 5 covers the foundations, including the policy that decides which tools are approved and who has oversight of them.

If these systems are weakest where knowledge is rarest, then knowing something uncommon very well counts for more now, not less. The people best placed are the ones with real depth in a narrow field who'll also learn how the tools behave, and that's the pairing I'd back. One without the other leaves you either unable to use them or unable to check them.

None of this asks you to track every release. It asks you to understand the board, which has not changed and is unlikely to.

Uglow's drawing game is still the clearest version of it. The machine has a picture of what a banana looks like, and it will keep asking you to draw its banana rather than yours.

Kristina Agustin is the Founder and Principal Digital Navigator of Southern Sky AI, helping maritime and professional organisations adopt AI with capability and good governance.

05

The Learn series

Five short episodes taking the full field briefing in order.

With thanks to Rob Howard at Innovating with AI, whose Plinko analogy this article borrows.

Read your position

Where does your organisation sit on the map right now?

The AI Baseline reads your current position in about five minutes and shows you the next plain move.

Get your baseline

Prefer weekly reading? Join the Chart Room Dispatch.