January 10, 2026
0 min
Sep 21, 2026
5:30 pm
-
8:30 pm
Parallax HQ, The Elbow Rooms, 64 Call Lane, Leeds, LS1 6DT
Jo Walton
Operations Manager
No items found.

Proof of value beats proof of concept: notes from our Applied AI with Anthropic

Jo Walton
Operations Manager

On Monday 21st September we sold out the room for the latest Applied AI, our community meet-up for people putting artificial intelligence to practical use.  It formed part of Leeds Digital Festival and was our first event with Anthropic. Every ticket sold went straight to UTC Leeds, raising £1000 to support the next generation of tech talent.

Is the proof of concept dead?

Sam Sutherland, Principal Tech Consultant at Parallax, opened with a problem we’re seeing on client work. A year or two ago, a prototype built in a few days would wow a room. Now everyone has seen someone build a 3D game from a couple of prompts, and plenty of client teams can knock up a demo themselves. A working prototype no longer proves much.  His answer is to stop proving AI can work, and start proving it creates value. Take invoice processing. The proof of concept asks whether you can extract the data and push it into internal systems. The proof of value asks whether anyone is actually doing less work, at an error rate the business can live with. If someone has to check every output by hand, the automation might not be saving anything.

Sam made the point with a story from before AI. Fresh out of university, he built a machine-learning scheduler for a rail and construction site-access firm, pitched as a replacement for their scheduling team. The demo worked, then the schedulers started asking questions… Greg picks his kids up on Thursdays. Bob has a dodgy knee and the site is on a hill. A pile of hotfixes later, the tool was unusable and the project was cancelled.

What worked was months spent physically sitting with the schedulers and building what they actually needed, which turned out to be a spreadsheet in a browser wired into the systems behind it. That became the foundation of a product Sam worked on for eight years, and it went on to plan close to 20,000 shifts a week for a single client.

His warning is that AI makes it easier to travel the wrong road faster. With today's tools, he'd have patched that first scheduler prompt after prompt without ever questioning the starting point. The first useful fix probably wasn't machine learning at all, just a script to import the spreadsheets people were retyping by hand.

Sam’s habits for building with AI: 

  • Measure first. Get a baseline, make one small change, then measure again. If the outcome didn't improve, stop or change direction.
  • Run evaluation suites everywhere, including on processes with no AI in them. Define the behaviour you expect and test what you'll rely on.
  • Make decisions inspectable. Show users the inputs, rules and evidence behind a result, so they can challenge the logic rather than just the answer.
  • Sit beside the user. Feedback buttons get ignored. Watching someone do the work tells you far more.
  • Build what you don't know yet. Spend the time AI saves on better questions and on problems you couldn't have tackled before.

Code generation is largely solved. Now what? 

James Hall, Parallax co-founder and Tech Director, then sat down with Ian Massingham, who leads Applied AI for startups across EMEA at Anthropic. The two go way back, to an AWS user group in our old office around 2014.

Ian's view is that code generation is now largely solved for the bulk of everyday software work. That doesn't make software easy - it means there will be far more of it, much of it generated by agents without anyone specifying it, and someone has to decide what stays. In regulated industries especially, unchecked proliferation is a real risk.

So what can't be handed to a model? Ian's answer was taste and a real grasp of business rules and policy. Anything built for more than one person still needs solid engineering and design fundamentals behind it. He also urged teams to pick use cases the next model release won't make trivial, and to write evals demanding enough to prove an advantage.

On shadow AI, James said the organisations doing best start from real pain points on the ground. They find out where teams are struggling and bring in people, consultancies like ours included, to sit beside them and listen. Picking use cases from the top because they look good in a PowerPoint rarely works. Ian added that the same controls apply whether you're 10 people or 10,000, starting with approved tools and a clear policy.

For anyone shaping an AI roadmap, his advice was to specify the goal, not the steps, and keep revisiting that goal. Ideas that needed frontier capability a year ago can become trivial with the next model, and teams that sit on an idea miss it. James suggested keeping decision records, so you can see what you decided a year ago and why, and go again.

What this means for the work we do

Both talks highlighted that when building gets cheap, the value moves to deciding what’s worth building and proving it works for the people who will use it. 

We do this every day at Parallax, and it’s how we approached a recent project for a client in a highly regulated sector. Their subject matter experts were piecing together information by hand across several systems. We started with a proof of value rather than a proof of concept, with Sam leading the engineering. The experts kept every decision, each AI recommendation was measured against their own calls, and value was proven before anything was productionised. The platform is now part of their day-to-day operation.

What’s next? 

The next Applied AI is already in the works. Follow us on Eventbrite to hear first. If you've got AI prototypes that haven't made it to production, we'd like to hear what's stopping them. Get in touch, or download our AI white paper.