Why I Threw Away My Video Harness When Opus 5.5 Arrived

James Phoenix
James Phoenix

I spent £700 on video analysis. A week later, Opus 5.5 came out, and I threw away my custom harness.

I had been making videos for an 11+ application and building a workflow to improve them. The harness wrapped generation in adversarial review with Gemini 3.1 Pro: render a video, analyse it, find problems, feed the criticism back, and try again. I was spending money on the feedback loop because I wanted better videos.

Then I tried doing the task directly with Opus 5.5. It was so much better at writing the code that produced the videos that it outperformed the workflow I had built around the earlier models. The extra machinery had stopped earning its place. I dropped the custom harness and used Opus 5.5 for the task.

Here is the same unit-conversion lesson before and after. I kept the narration the same, so I can compare the visual explanation directly.

Earlier custom harness
Opus 5.5 rebuild

At 0:45, the earlier version connects small cards with an arrow. The rebuild shows 2 metres and 200 centimetres on the same scaled ruler. At 1:55, it places 3.4 kilometres on aligned kilometre and metre scales. I get a larger, more concrete explanation of why the conversion works, rather than just an instruction to multiply.

That is an uncomfortable thing to do after spending money and effort on a system. I had worked out how to make the old approach better. I had something I could explain, refine and keep improving. Throwing it away meant accepting that the best next step was to stop improving it.

The money I had already spent could not make the old workflow the right choice for the next video.

I want to be willing to discard an approach as soon as a better one proves itself on my actual work. A harness is built around a particular model’s capabilities and weaknesses. When those change, I need to reconsider the code that compensates for them. Something that was useful last week can become unnecessary this week.

This has also made me think harder about how much harness I want in the first place.

A minimal harness gives the model more room to choose its approach. I can describe the result I want and let it decide how to get there. That leaves more space for a stronger model to surprise me, including by finding an approach I would never have written into a pipeline.

Leanpub Book

Read The Meta-Engineer

A practical book on building autonomous AI systems with Claude Code, context engineering, verification loops, and production harnesses.

Continuously updated
Claude Code + agentic systems
View Book

It also leaves more room for variation. Different runs can take different routes, make different creative choices, and need different amounts of correction. I gain freedom and accept a less predictable process.

A heavier harness can make that process more consistent. Fixed stages, templates, restricted choices and enforced checks reduce the space the model can move through. That can be valuable when I need repeatable output. But I can also restrict it to the approaches I knew how to encode. A stronger model may be capable of better work while my workflow keeps steering it through the same narrow route.

I can make the process more predictable while limiting the quality of the result.

I am increasingly interested in keeping the requirements for the finished work clear while giving the model freedom over the method. For a teaching video, the explanation still needs to be correct, the text readable, and the visuals timed sensibly. I can inspect those properties without prescribing every creative decision in advance.

The balance depends on what variation costs me. I am happy to explore different approaches when making a video. I would want tighter controls for a repetitive task where an unexpected choice creates expensive downstream work.

For this task, Opus 5.5 changed that balance enough to make my custom harness obsolete. I want to keep what I learned from building it and stay willing to throw away the code.

Explore it visually

Topics
Agent ArchitectureAi AgentsWorkflowsPrompt Engineering

Newsletter

Become a better AI engineer

Weekly deep dives on production AI systems, context engineering, and the patterns that compound. No fluff, no tutorials. Just what works.

Join 306K+ developers. No spam. Unsubscribe anytime.


More Insights

Cover Image for Build Features in the CLI Before the API

Build Features in the CLI Before the API

Make the next app as much of a CLI app as possible. Feature work happens in the terminal, in-process, against the real development database. The HTTP API is a thin adapter you add afterwards, and a lint rule fails CI if it ever gets a route the CLI does not have.

James Phoenix
James Phoenix
Cover Image for Useful Output Across Three Modes of Work

Useful Output Across Three Modes of Work

I compare typing, interactive AI assistance, and automated loops by useful output per hour of my attention, including review, compute, and failure costs.

James Phoenix
James Phoenix