Yesterday, OpenAI released a large batch of new mathematical results produced by an unreleased internal model.
Some of them go after problems the best mathematicians in the world have worked on for decades. The community is still verifying the results, but if even part of it holds up, it is a big moment.
When you read developments like this, it’s easy to forget that it’s still less than four years since ChatGPT came out (November 2022). That blows my mind, because we’ve seen so much change in such a short period of time.
Listen to a legal innovation or tech podcast from before November 2022, and chances are you won’t even hear the term “AI” at all. Yet here we are with announcements like this almost every day, multi-billion dollar funding rounds, trillion-dollar IPOs, AI-native law firm launches - the list goes on.
Our industry has undergone more change in four years than at any time in its history. But with so much change in such a short period, it’s easy to lose perspective. Without that perspective, we forget that:
AI was pretty bad just four years ago
We’ve seen incredible change in four years
We’re still right at the start
Things will probably be unrecognisable four years from now.
So several weekends back, to help provide some perspective (and frankly just out of curiosity), I ended up building The AI Time Machine, a working museum of AI so you can go back and chat with any AI model from the last 60 years.
I hope you have fun testing it out, remember how bad AI was until recently, and get a fresh perspective on where things may be going next.
How it works
It is a working museum of conversational AI. You get a familiar chat box, and a model picker that goes all the way back to 1966.
You can ask ELIZA, Joseph Weizenbaum’s famous MIT chatbot, for advice and watch it turn your words back into questions. You can probe PARRY, Kenneth Colby’s simulated paranoid patient, and watch it go from guarded to hostile. You can watch GPT-2 start a sentence fluently and then lose the plot. You can get emoji-bombed by GPT-4o. And then you can put exactly the same question to a current frontier model. (I also included Jev - though it’s not a chatbot - for comparison.)
A quick word on how it works:
The old bots are the real thing. ELIZA, PARRY, RACTER, Dr. Sbaitso, SmarterChild and friends were effectively deterministic rules, not neural networks. So the museum re-implements their original algorithms, and they run in your browser.
The early transformers are mimics. GPT-1 and GPT-2 are no longer available from anyone, so the museum mimics their documented failure modes (e.g., GPT-2’s tendency to collapse into repetition).
Retired large language models are recreations. GPT-3, the original ChatGPT, early Claude and many others no longer exist on any API. Providers retire old models. So, a current model plays each one under a brief based on its model card and what is available in the public record: what it could do, how it formatted answers, how often it made things up, its signature quirks, etc.
Current models answer as themselves. If you bring your own API key, today’s Claude and GPT models answer live.
It is a hobby project, built with AI, and the code is open source, so it’s best to treat it as more of an educational exhibit than a primary source.
Also, models increasingly do far more than just chat. They write code, create images, make videos, spawn sub-agents, and so on. For now, the museum is focused simply on how they talk. Maybe in some future update I’ll look to capture some of these other functionalities.
How bad was it, really?
It is easy to forget. So here is a quick reminder of what the launch version of ChatGPT was like in late 2022:
It had a tiny memory. The underlying model could hold roughly 4,000 tokens, or about 3,000 words, across your prompt and its answer. Forget a data room. It could not hold a decent-length contract.
It was frozen in time. Its knowledge stopped in 2021 and it could not search the web. Ask about anything recent and you got a confident guess or a polite refusal.
It could not do sums. Simple multiplication regularly went wrong. The idea that a descendant would publish hundreds of maths papers four years later would have sounded like science fiction.
It only talked. No documents, no tools, no agents. You typed, it typed back, and you copied and pasted.
I remember showing early ChatGPT to lawyers and friends in 2023. The reactions were split between “this is magic” and “this might work for other industries, but ours is different”.
Then things moved quickly. GPT-4 arrived in March 2023 and OpenAI said it scored around the 90th percentile on a simulated bar exam. Context windows grew from pages to whole books. Models learned to read PDFs, search the web, write and run code, and use tools. “Reasoning” models arrived that think before they answer. Then agents that can work for hours on a task.
And now, four years later, we are arguing about whether AI-generated proofs of open mathematical problems meet the publication standards of the Institute for Advanced Study!
How do we know the recreations are any good?
If you ask a modern model to “pretend to be GPT-3”, it will happily do so. The problem is that it tends to be too good. It slips in vocabulary, knowledge and skills that GPT-3 probably didn’t have. And I’ve found that it becomes a bit of a caricature.
So the Time Machine has an automated test harness that checks the exhibits against the historical record. It has three layers:
Engine checks. Simple, repeatable tests. Does ELIZA’s “can you” rule trigger? Does PARRY slide from guarded to hostile when you push it? Does GPT-2 fall into its repetition loops?
Historical ground truth. Tests written from primary sources (Weizenbaum’s 1966 paper, Colby’s PARRY transcripts, the model cards), separately from the code. These exist to catch the code faithfully doing the wrong thing.
An AI judge. A current frontier model scores each recreation from 0 to 10 against a brief for that model, looking for anachronisms: words that postdate the model, abilities beyond what it could do at the time, and missing signature quirks.
Why am I sharing this? Who cares?
Fair question. Partly because a museum of old chatbots is fun and I couldn’t find another one online. But also I think this might change our perspective a bit.
1. Stop assuming the AI you see today is the AI you will have tomorrow
Most AI strategies I see are built on a snapshot. A firm runs a pilot, finds that the tool is good at X and poor at Y, and writes a policy and a business case around that. Eighteen months later, Y is solved and the business case is out of date.
If you had written your strategy in early 2023 based on launch ChatGPT, you would have concluded that AI cannot read contracts, cannot do research and cannot be trusted with citations. Two of those three are no longer true, and the third has moved a long way. Plan for the capability curve, rather than where we are today.
2. Be careful what you write off
Many lawyers formed their view of AI in 2023, often from a bad experience or headline, and some I speak to have not updated it since. That is understandable. But a judgment formed on a model that is now a museum piece is a judgment about the past. The person in your firm who says “I tried it, it was rubbish” should be asked: which model, and when?
3. The business model will lag, so start now
Business model change always lags technology change. The technology has moved from “a cool toy that makes stuff up” to “new results in mathematics” in four years. Yet most firms still price, staff and train much as they did in 2022. The gap between what the technology can do and how firms are organised to use it is getting wider.
I think this is where both the risk and the opportunity sit. See, for example, my AI Firm Index project, which tracks new law firm launches. It has more than doubled in just a few months and I’m seeing new firms apply almost every day.
It’s become a bit of a cliche to say that “the AI you are using today is the worst AI you will ever use”. We nod, then carry on planning and building on the capabilities that exist today, rather than preparing for the capabilities that will exist tomorrow. I hope this brings some fresh perspective.






Excellent. A very interesting and informative post, likely to be didactic to a wide audience.