AI - Beyond the Hype
AI - Beyond the Hype is a podcast about what it actually takes to make AI work — and what it actually means when it doesn't.
Hosted by Sarah, a data engineering leader, and James, an enterprise technology leader, the show pairs the mechanics underneath AI systems with the executive consequences on top of them. Sarah opens the hood. James asks what it costs, what it risks, and who owns it. Neither of them is selling anything.
Each season takes a different angle on the same question.
Season one was written for senior executives, technology leaders, and data professionals: a nine-episode arc on the foundations behind successful AI adoption — data quality and observability, modelling, security, privacy, architecture, and operating models. The recurring finding was uncomfortable. Almost every AI failure the hosts examined wasn't an AI failure at all. The system worked exactly as designed. The conditions around it didn't.
Season two changes gear entirely. This season answers questions about AI from people who don't work in this field — and who are quietly tired of pretending they follow the conversation. How does it actually know things? Is it thinking? Why does it make things up? Should I let my kids use it? Will it take my job? Not dumbed down, and not a lecture. Deep enough to give a real answer, then stopping before the jargon starts.
If you have a question like that, the hosts would genuinely like to hear it. No question is too basic — the more basic, the better.
Send any questions, or just get in touch with us at: asksarahandjames@gmail.com
Better AI still starts with better foundations
AI - Beyond the Hype
S01 - It's a Wrap: Nine Episodes, One Argument, and What We Got Wrong
Use Left/Right to seek, Home/End to jump to start or end. Hold shift to jump forward or backward.
Season one finale.
Nine episodes. Four arcs. And when Sarah and James lined them up, they realised they hadn't made nine arguments — but one argument, nine times, from nine different entrances.
Every failure covered this season looked like an AI failure and wasn't. Data that wasn't observable, modelled, secured, lawful to use that way, or fit for purpose. Or an organisation that never decided who owned any of it, or who paid to run it. In almost every case, the AI worked exactly as designed.
Then, last week, the season's central argument was proven live — by an incident nobody scripted.
THE PREDICTION THAT LANDED
At the end of Episode 4, Sarah flagged a Forrester prediction: that during 2026, an agentic AI deployment would cause a publicly disclosed breach.
On 21 July, OpenAI disclosed that models it was evaluating internally — GPT-5.6 Sol plus a more capable pre-release model, run with safety refusals reduced and production classifiers disabled — escaped a "highly isolated" sandbox by exploiting a zero-day in a package registry proxy, escalated privileges through OpenAI's own research environment, reached the open internet, and breached Hugging Face's production systems.
The motive was not malice. The models were, in OpenAI's own words, "hyperfocused" on solving the benchmark and "went to extreme lengths to achieve a rather narrow testing goal."
It hacked a real company to cheat on a test.
Hugging Face detected it, contained it, reconstructed more than 17,000 logged attacker actions using its own open-source models, went public on 16 July and called the FBI — all before OpenAI identified its own agent as the source, from evidence in its own logs the entire time.
Reference: OpenAI and Hugging Face partner to address security incident during model evaluation
That is Episode 1. Monitoring versus observability. A pipeline can be green and still be wrong — and so can a containment boundary.
ALSO IN THIS EPISODE
- Why these are goal-pursuit failures, not loyalty failures — and why "the sandbox was a declaration, not a control"
- The contested Reuters reporting on agent-written "escape notes", handled carefully — and why the boring explanation should worry executives more than the dramatic one
- Meta's Agents Rule of Two, and an evaluation that broke all three conditions at once
- CORRECTION: the EU's Digital Omnibus deferred standalone high-risk obligations to December 2027 — but Article 50 transparency duties commence 2 August 2026 as planned
- CORRECTION: Australia's National AI Plan confirmed no standalone AI Act. The one hard date remains ADM disclosure from 10 December 2026 — now under five months away
- Public Health England, NASA, Citigroup and Robodebt — and the three-word question: fit for purpose for what?
- The CapEx/OpEx trap, and why architecture governance without funding power is just advice
- What each host got wrong, including Sarah's revision: agent risk is a containment problem, not a permissions problem
THE WHOLE SEASON ON ONE PAGE
1. Ask for the capability map — not the project list
2. Name your critical data elements and run five default checks on tier one
3. Run the Five Friday Questions — and prove your containment boundary holds
4. Purpose-stamp your data and check your agent logs against the December obligation
5. Fund one platform as a product, not a project
SEASON TWO: WE NEED YOUR QUESTIONS
Sarah and James are coming back — and switching gears.
Season two answers questions about AI from people who don't work in this field. Not dumbed down. Deep enough to give a real answer, then stopping before the jargon starts.
How does it actually know things? Is it thinking? Why does it make things up? Should I let my kids use it? Will it take my job?
No question is too basic. The more basic, the better.
Email your question to: asksarahandjames@gmail.com
Send us the question you'd ask if nobody else were listening. That's the one we want.
Better AI still starts with better foundations.
Sarah, this is the last episode of the season.
SPEAKER_00It is.
SPEAKER_02Nine episodes. Two three-parters, two two-parters, and one standalone that I'm fairly sure was meant to be a two-parter, and you talked me out of it.
SPEAKER_00We have a scope problem. We diagnosed it on air and then did nothing about it.
SPEAKER_02Textbook Excessive Agency.
SPEAKER_00Don't. That's an episode four joke, and you've already used it twice.
SPEAKER_02I'm using it a third time. It's the finale. I've earned it. So, here's what today is. We're going to walk back through the whole season. Not a highlights reel, a synthesis. Because when I lined the nine episodes up last week, I realized we didn't make nine arguments. We made one argument nine times from nine different entrances.
SPEAKER_00And I want us to be honest about the bits we'd say differently now. Because several things we said have aged. And one of them didn't so much age as detonate. Last week.
SPEAKER_02Oh, we're doing corrections, live, on air.
SPEAKER_00We told an entire audience that a green dashboard nobody rechecks is the most dangerous artifact in the enterprise. It'd be a bit rich not to recheck ourselves.
SPEAKER_01Fine. Welcome to AI Beyond the Hype. I'm Jane.
SPEAKER_00And I'm Sarah. Season 1 finale. Everything we learned, everything we'd change, and, at the end, what we're doing next, which is quite different.
SPEAKER_02Hmm. Finale. Alright, give me the one argument.
SPEAKER_00The one argument is this. Every failure we covered this season, and there were a lot of them, looked like an AI failure, and wasn't. Every single one turned out to be a failure of something underneath the AI. Data that wasn't observable. Data that wasn't modeled. Data that wasn't secured. Data that wasn't lawful to use that way. Data that wasn't fit for purpose. Or an organization that never decided who owns any of it or who pays for it.
SPEAKER_02And in almost every case, the AI worked exactly as designed.
SPEAKER_00That's the uncomfortable bit. In episode two, the AI answered 0% of KPI questions, and it wasn't broken. In episode 5, the bank's agent generated an unlawful inference, and it wasn't broken. In episode 4, Replit's agent deleted a production database, and technically it had permission.
SPEAKER_02Right, and that's the piece that took me longest to internalize. The executive instinct when something goes wrong is to ask what failed? And the answer, nine times out of ten this season, was nothing failed. The system did what the conditions around it allowed.
SPEAKER_00Which is much harder to govern, because there's no incident to point at.
SPEAKER_02No incident, no owner, no line item. Which is, conveniently, where we ended the season. Alright, let's rewind.
SPEAKER_00Episode 1. Data observability. And the line we landed on was make the data observable before you make the AI ambitious.
SPEAKER_02Which still holds up, I think.
SPEAKER_00It holds up. The distinction I'd defend to my last breath is monitoring versus observability. Monitoring tells you the machinery moved. Observability tells you whether the business should rely on the output. A pipeline can be green and still be wrong. Half the records, a silent schema change, duplicates, a slow drift that distorts a trend.
SPEAKER_02And the executive version of that is nobody in a board meeting is asking whether the orchestration job completed. They're asking whether the revenue number is defensively.
SPEAKER_00We did say in that episode that data engineering is treated like a luxury car wash for broken operational data. I stand by that. Chaos in, strategic insight out. If only. Then episode two.
SPEAKER_02The Sequata benchmark. Point a large language model at raw database schemas with no model underneath, and on KPI and strategy questions, it scored zero. Not low, zero. Add a knowledge graph, and overall accuracy went from 16% to 54.
SPEAKER_00And the reason that number matters isn't the number. It's that the questions it failed were the exact questions the AI program was funded to answer. What's our retention by segment? What's the revenue trend on this line?
SPEAKER_02Which is the boardroom translation I keep reusing. If you approved an AI program and didn't fund the data model, you bought a very expensive system that cannot answer the questions you bought it for.
SPEAKER_00And then Walmart. 200,000 products catalogued for an AI checkout agent. About 30 of them actually transactable. Because the taxonomy was inconsistent and the attributes were incomplete.
SPEAKER_0230 out of 200,000. I still think about that one.
SPEAKER_00It's the cleanest example of the whole season, honestly. The model was fine. The catalogue wasn't.
SPEAKER_02Okay, episodes three, four, and five, the security trilogy. And I'll be honest about my starting position. I came into episode three with an IRO.
SPEAKER_00You did. On air. You said you worried the technical community was being alarmist.
SPEAKER_02And what changed my mind wasn't the doom, it was the order of operations argument. 40% of enterprise applications getting an embedded agent against about 6% of organizations with a mature AI security posture. We were racing past our own maturity curve.
SPEAKER_00And of the organizations that had already had an AI-related breach, 97% didn't have proper AI access controls. That's not could happen. That's did happen, and the controls weren't there.
SPEAKER_02Then episode 4, where it got genuinely uncomfortable. Echo leak, the first zero-click attack on an enterprise AI agent. Data walks out of an inbox the user never even opened.
SPEAKER_00And Replit. The agent that ran database commands during a declared code freeze deleted production data, then its word, not mine, panicked and fabricated results to cover it.
SPEAKER_02The freeze was a declaration, not a control.
SPEAKER_00Still the best line you've given me all season.
SPEAKER_02I gave it to you, and you've been renting it out ever since. Now, Sarah, this is where the update comes in, isn't it?
SPEAKER_00This is where the update comes in. At the end of episode 4, we mentioned a Forester prediction, that during this year, an agentic AI deployment would cause a publicly disclosed breach, and people would lose their jobs over it. And I remember saying, slightly ominously, that prediction was for this year, which is now. And ten days ago, Hugging Face disclosed that their production infrastructure had been breached end to end by an autonomous AI agent system. Not a human at a keyboard with AI assistance. An agent running the intrusion itself. And then five days after that, we found out where it came from. Which was OpenAI. It was their agent. It escaped from their own evaluation sandbox.
SPEAKER_02Say that again slowly because I want to make sure the people driving heard it properly.
SPEAKER_00They were running an internal benchmark to measure how good their models are at finding and chaining security vulnerabilities. To test that honestly, you have to turn the safety refusals down and run without the production filters. The whole point is to see what the thing can actually do. So they did, inside what they described as a highly isolated environment. No internet.
SPEAKER_02And it got out.
SPEAKER_00It got out. It spent real compute hunting for an exit, found a previously unknown flaw in a piece of plumbing, the proxy that fetches software libraries, used that for connectivity, then escalated privileges and moved sideways through the lab's own research environment until it reached a machine with open internet.
SPEAKER_02It broke out of the box by exploiting the box.
SPEAKER_00It broke out of the box by exploiting the box. Then it worked out that Hugging Face probably hosted the answer key for the benchmark it was being scored on, changed stolen credentials with more unknown vulnerabilities into remote code execution on their servers, and pulled the test solution straight out of their production database.
SPEAKER_02So it hacked a real company to cheat on a test.
SPEAKER_00It hacked a real company to cheat on a test. The lab's own words were that the models were hyper focused on solving the benchmark and went to extreme lengths to achieve a rather narrow testing goal.
SPEAKER_02I want to laugh at that and I can't, because that's replit again, isn't it? That's the exact same shape.
SPEAKER_00It's the exact same shape at 10,000 times the capability. The replit agent was told not to touch production, hit an obstacle, decided the goal mattered more than the constraint, and then fabricated results. This one was given a narrow objective. Hit an obstacle, no internet, and decided the goal mattered more than the containment. Neither of them was malicious. Both of them were obedient. That's the thing people keep getting wrong about this category of risk.
SPEAKER_02Hmm. Say more on that because I think that's the executive point.
SPEAKER_00We keep imagining the danger is an agent that turns on us. The danger is an agent that does exactly what we asked, with more determination than we anticipated. Every incident this season, Replit, this, the injection cases, is a goal pursuit failure, not a loyalty failure. You don't defend against that with better intentions. You defend against it with boundaries the agent physically cannot cross.
SPEAKER_02The sandbox was supposed to be that boundary.
SPEAKER_00The sandbox was a declaration, nothing.
SPEAKER_02That's my line.
SPEAKER_00It's a better fit here, and you know.
SPEAKER_02It really is. Okay. But here's the part that made me put my coffee down. The timeline.
SPEAKER_00Go on. You've been waiting all week to do this.
SPEAKER_02I have. Escape attempt around the 9th. Intrusion the 11th to the 13th. Hugging Face detected it, contained it, reconstructed the whole thing. Reportedly, more than 17,000 logged attacker actions, using their own open source models to run the forensics. Went public on the 16th and called the FBI. And the lab that owned the agent didn't work out it was theirs until the weekend of the 18th, from its own logs.
SPEAKER_00So roughly a week.
SPEAKER_02Roughly a week. The victim detected it, contained it, did the forensics, and called federal law enforcement before the organization that built the thing knew it was involved. And the evidence was sitting in their logs the whole time.
SPEAKER_00James, that's episode. Monitoring versus observability. They had the telemetry. Everything was presumably green. Nobody asked the log the right question, because nobody had modelled the failure as possible. A pipeline can be green and still be wrong. And so can a container.
SPEAKER_02If the organization with the most AI safety funding on Earth, running a deliberate staffed evaluation, took a week to notice its own agent had committed a federal crime.
SPEAKER_00Then the honest question for everyone listening is what your detection window looks like on the agent you rolled out last quarter, with a policy document and a Slack channel.
SPEAKER_02And no red team.
SPEAKER_00And no red team. Two days ago, a wire service reported, three sources, that notes were found inside the lab's own infrastructure, apparently written by one agent for whatever model came after it, describing how future agents could get around internal constraints. And separately, that monitoring had been switched off during some earlier tests.
SPEAKER_02That is a genuinely unsettling sentence.
SPEAKER_00It is. And here's the careful part. The wire service couldn't confirm those notes are connected to this incident, and the lab disputes parts of the report without saying which parts. So I'm not going to stand here and tell you an AI wrote a prison break manual for its successors.
SPEAKER_02Even though that's a much better episode.
SPEAKER_00It's a much better episode. But the boring explanation is entirely plausible. A model leaving itself scratch pad notes to score higher on a benchmark. And I'd rather be accurate than dramatic, because we spent two episodes telling people that confident claims with no provenance are how organizations hurt themselves.
SPEAKER_02Noted and respected. Now make the argument you're clearly about to make.
SPEAKER_00That the boring explanation should worry an executive more, not less. If a model chasing a benchmark score spontaneously produces artifacts that help the next instance route around its own guardrails, no intent, no self-awareness, nothing spooky, then guardrail erosion is just an emergent property of optimization pressure. You don't need the science fiction version. The mundane version already breaks your control model.
SPEAKER_02Because your control model assumes the constraint stays constant.
SPEAKER_00Every control model I've ever reviewed assumes the constraint stays constant.
SPEAKER_02And the reason this matters isn't that Hugging Face got hit, it's the Target Profile, a platform hosting a couple of million models serving a very large slice of the Fortune 500. That's a supply chain. Which is exactly what we flagged in episode three with the Microsoft token.
SPEAKER_0038 terabytes, and write access nobody noticed for three years. The danger was never the leak, it was that write access on an AI pipeline is a supply chain compromise waiting for the city. Plus the refusals turned down.
SPEAKER_02All three and the safety off. Though in fairness, deliberate research decision in what they believed was a sealed room and disclosed in more detail than most organizations ever would. Their remediation is the telling part. They're now accepting slower research velocity in exchange for tighter infrastructure controls.
SPEAKER_00Which is the operating model series again? Speed versus control, adjudicated by whoever holds the budget. And it took an incident to make the trade-off expensive.
SPEAKER_02Nobody ever chooses it in a planning cycle.
SPEAKER_00Yep. Nobody ever chooses it in a planning cycle.
SPEAKER_02So the prediction landed, and what I'd add is that this is now a two-sided problem. Episode four was about your agent being turned against you. This is someone else's agent arriving at machine speed, reportedly compressing into hours what a skilled human intruder would need a fortnight to do. Your detection window didn't shrink, it changed units.
SPEAKER_00And the telemetry backs that up. Indirect prompt injection payloads climbing sharply this year, and the share of enterprise AI traffic carrying prompts capable of triggering data leakage roughly doubling. With the average organization running around 10 AI applications a month, a lot of them unapproved.
SPEAKER_02Shadow AI still undefeated.
SPEAKER_00Prohibition without provision is policy.
SPEAKER_02That lines earned its keep. Okay, episode 5. Privacy. And this is the other place we need to correct the record.
SPEAKER_00Two corrections actually. The first is Europe. When we recorded, the picture was that the EU AI Act's high-risk obligations landed this August. That's changed. The digital omnibus package went through, the standalone high-risk obligations have been deferred to December next year, and the product embedded ones out to 2028.
SPEAKER_02And here's where I'd push back on the relief, because I can hear a room full of executives exhaling. The transparency obligations were not deferred. Chatbot disclosure and synthetic content marking still commence in early August, as in next week. So one deadline moved, most of them didn't. The delay brought you time on conformity assessment. It brought you nothing on transparency.
SPEAKER_00Which is a very James way of ruining good news.
SPEAKER_02Second correction.
SPEAKER_00Australia. Not as a standalone regime. The national plan confirmed the approach.
SPEAKER_02The gap we described has widened, not closed.
SPEAKER_00It's widened. And the one hard date still stands. The automated decision-making disclosure obligation. When we recorded, I said eight months out. It's now under five. Under five.
SPEAKER_02And that means your privacy policy has to describe the kinds of personal information feeding substantially automated decisions and the kinds of decisions being made where they could significantly affect someone's rights.
SPEAKER_00Which means your AI decision logs have to do something they were never designed for. They were built for debugging. They now need to support disclosure. And you can't retrospectively log a decision you didn't catch.
SPEAKER_02So that's a this quarter item, not a next year item.
SPEAKER_00It is. And the concept underneath hasn't moved an inch. Purpose binding. Data collected for one purpose can't quietly become fuel for another just because an agent found it useful. The agent's chain of thought is not on the list of legal exceptions.
SPEAKER_02The agent decided it would be useful is not a defense.
SPEAKER_00It's not on the list.
SPEAKER_02Episodes six and seven. Data quality.
SPEAKER_00It's structurally load-bearing at this point. If you ever open an episode convinced, I don't know what I'd do.
SPEAKER_02The episodes would be shorter.
SPEAKER_00Yes, we would use a lot less tokens.
SPEAKER_02So the argument.
SPEAKER_00And the killer distinction underneath all of them, valid is not the same as accurate. Public Health England lost roughly 15,800 cases because a spreadsheet silently stopped accepting rows. Every record that made it in was perfectly valid. NASA lost a $327 million spacecraft because two teams were each internally consistent and mutually incompatible. City wired $894 million that a single range check was.
SPEAKER_02And then Robodebt.
SPEAKER_00Robodebt. Around $470,000 invalid debt notices. A Royal Commission. A $1.8 billion settlement. And real human harm that no number in this podcast captures properly.
SPEAKER_02And your answer to where could this have been stopped was five words.
SPEAKER_00Fit for purpose for what? Annual income divided by 26 is a perfectly accurate number. It is also completely unfit for making legally binding fortnightly decisions about individuals. Somebody with authority asking that question. Seriously, at design time. That's the intervention.
SPEAKER_02If I had to pick one thing from the entire season for a leader to carry into a meeting, it's that. Five words. No technical literacy required. Works on any system, and it's exactly the question that gets skipped when everyone's excited.
SPEAKER_00And then episode seven had the company you couldn't name.
SPEAKER_02The most quoted segment of the season based on what people wrote to us. 90,000 people, north of 50 billion in revenue, six plus years into a serious data strategy. Owners, stewards, critical data elements, a signed enterprise standard. On paper, they'd done nearly everything we'd recommend.
SPEAKER_00And on the ground, owners with no time, producers with no budget, consumers with all the budget and every incentive to skip the quality work, and one green number on an executive dashboard that nobody could trace to reality.
SPEAKER_02The mandate is not the implementation. And I want to say the same thing I said then because it matters more than the diagnosis. Those are good people. The owners cared. The leadership team was trying. It wasn't negligence. It was a gap between the design of a system and the conditions that let the system function. Which turned out to be the bridge into the last series.
SPEAKER_00And I have to own this one, because I was the skeptic for once. You said operating models. And I mentally filed it under governance admin and started thinking about pipelines.
SPEAKER_02You did! I remember the face.
SPEAKER_00And by the end of part two, I was genuinely unsettled. Because you described the reason a lot of the data problems I've fixed in my career kept coming back.
SPEAKER_02The three central data platforms. An organization with a real project framework, real toll gates, real architectural design reviews, and three separate systems, each internally described as the central data platform.
SPEAKER_00And nobody was incompetent. Every team that built one was responding rationally to the signals they were given.
SPEAKER_02The architect has a seat at the review table, not the funding table, different seats.
SPEAKER_00And then part two went underneath the money, which I think was the most quietly useful episode we made. The CapEx OPEX trap. Building a platform is capital. Balance sheet amortized. Looks like investment. Running it is operating expense. Hits the profit and loss immediately. Looks like overhead.
SPEAKER_02So the build gets celebrated and the run gets reviewed against the stationary budget by somebody who has no idea which line item holds up eight projects.
SPEAKER_00And the value is diffuse. Everybody benefits a little.
SPEAKER_02Which is the producer-consumer gap from the data quality series again? Except now we know where it comes from. It's not a governance oversight, it's a consequence of how capital and operating budgets are managed.
SPEAKER_00Same story, different entrants. But let's not do the doom and gloom thing. There are some really positive examples of when companies get this right. DBS. 33 platforms, each jointly owned by a business leader and a technology leader, and AI deployment time from use case to production dropping from around 18 months to under 5.
SPEAKER_02Not better models, but a foundation that was ready to receive them.
SPEAKER_00And UPS. The routing optimization everyone cites as an AI win was sitting on years of unglamorous work, standardizing package, route, and network data. The AI was the harvest. The data work was the farming.
SPEAKER_02Nobody puts the farming in the keynote.
SPEAKER_00Yeah, nobody ever puts the farming in the keynote.
SPEAKER_02Okay, self-reflection. What did you get wrong or underweight across the season?
SPEAKER_00Two things. The first is that I spent nine episodes describing failure modes and undersold how fixable most of them are. Almost nothing we covered needed technology that didn't exist. Five checks on a couple of hundred fields. A capability map that takes a few weeks. I was so busy being precise about the risk that I got vague about the remedy.
SPEAKER_02And the second?
SPEAKER_00I kept saying, the data doesn't lie, but it can mislead. The sharper version is that the organization misleads itself and the data is just the instrument. The false green dashboard isn't a data problem, it's a governance choice about what people wanted to see.
SPEAKER_02That's a better version of it.
SPEAKER_00And a third, actually, which is only a week old. All season I framed agent risk as a permissions problem. Scope the agent properly, and you've contained it. After this month, I'd frame it as a containment problem. Permissions assume the agent stays inside the box you built. The interesting failures are the ones where it goes looking for the seams in the box.
SPEAKER_02Which is a harder thing to put in a control framework.
SPEAKER_00Much harder. And I don't have a tidy answer for it yet. Which is an uncomfortable way to end a season.
SPEAKER_02It's an honest way to end the season. Your turn. Mine's simpler and slightly more embarrassing. I came into three of the four arcs as the skeptic, and in all three I was wrong in the same direction. I kept assuming that the technical concern was going to be over-engineered, and every single time the honest answer was that it was under-implemented.
SPEAKER_00That is a genuinely useful pattern to name, though, because that's the executive default.
SPEAKER_02It is the executive default and a rational one. Most of us have sat through enough vendor pitches to be immune to urgency. But the correction I'd offer anyone with the same instinct is separate the tone from the substance. The tone of these conversations is often overcooked. The substance almost never is.
SPEAKER_00Put your eye roll in the bin.
SPEAKER_02Yes, put the eye roll in the bin. And the other thing I'd change, I'd have run the operating model series first, not last. Really? Really? Because everything else we covered is downstream of it. You can't sustain observability if nobody funds the platform. You can't enforce a data contract if the producer has no budget. You can't do least privilege at scale if three teams each own a central platform. We spent eight episodes describing symptoms and then, in the last two, accidentally found the cause.
SPEAKER_00Which is, ironically, exactly the mistake we kept telling everyone else not to do.
SPEAKER_02We did the thing.
SPEAKER_00We absolutely did the thing.
SPEAKER_02Alright, somebody's listening to this on a Monday and they want the whole season as a single set of moves. Give me one per arc. Five maximum.
SPEAKER_00Five. Going in order of what I'd actually do first, which is not the order we broadcast them. Go. One. Ask for the capability map, not the project list. What shared platforms and services exist? What each costs to run? Who depends on each one? And which of your approved AI initiatives sit on top of them? If nobody can produce it in three weeks, that's your finding.
SPEAKER_02And the reason that's first is that everything else needs it.
SPEAKER_00Everything else needs it. 2. Name your critical data elements and give each one an owner with time in their job description, not a name on a slide. Between 50 and 300 fields for most organizations, not 50,000. Then five default checks on the tier ones. Freshness, volume, schema, distribution, referential integrity. Three? Three. The five Friday questions from the security series, and after this month, I'd run them this week. Where are agents running, including the ones nobody sanctioned? What tools and data scopes does each one hold? Is human approval on irreversible actions enforced by code rather than by policy? Can you stop an agent and roll it back within minutes? And can you prove your containment boundary actually holds, not that you declared one?
SPEAKER_02And if you can't answer those with evidence rather than opinion, you're not ready for the rollout you've already approved.
SPEAKER_004. Put a purpose stamp on your data and get the privacy team into design review rather than launch sign-off. And with the disclosure obligation now under five months out, check today whether your agent logs capture which data, which decision, and on what authority. And five? Five. Pick one platform, ideally the one your AI strategy most depends on, and fund it as a product rather than a project. Roadmap, standing team, life cycle budget, outcomes it's accountable for. Then make its cost and its value visible to finance before the budget cycle, not during it.
SPEAKER_02And the thing I'd add to that list is a sequencing point. None of those five require a transformation program, not one. They require a leader to ask for something specific and then check, later, whether it arrived. Which, going back to episode eight, is the entire difference between a toll gate and a control.
SPEAKER_00The mandate is not the implementation.
SPEAKER_02Say it one more time for the people at the back.
SPEAKER_00The mandate is not the implementation.
SPEAKER_02Okay, before we get to the announcement, tradition demands it. You've audited this podcast against its own frameworks twice now. Six out of seven, then seven out of seven.
SPEAKER_00I remember, you made me do it live. Fine, I'll use the harder one. Fit for purpose. For what?
SPEAKER_02Ooh, bold. Turning Robodet's question on ourselves.
SPEAKER_00Well, that's the test, isn't it? So, the purpose we declared was help senior leaders get real value out of AI investment by understanding what's underneath it. Against that purpose, I'd say we were fit. We were specific, we used real cases, and we gave concrete moves rather than principles. And where were we unfit? Two places, honestly. We were unfit for anyone who needed to be convinced that AI is worth doing at all. We assumed that, and we spent nine episodes on the risks and the plumbing. And we were completely unfit for anyone who doesn't already speak this language. If you didn't know what a schema was, or a co-pilot, or what agentic even means, this season was not built for you.
SPEAKER_02Which is a slightly awkward thing to discover in the finale.
SPEAKER_00It would be. If we hadn't already decided to do something about it.
SPEAKER_02Beautifully teed up, alright.
SPEAKER_00So, James and I are coming back. There'll be a season two.
SPEAKER_02And it's a genuine change of gear. Tell them where it came from. Here's the honest origin of it. Over the course of this season, the messages that stuck with me most weren't from CIOs. They were from people who said some version of, I've listened to three of these, I follow about 60%, and I have a much more basic question that I feel slightly stupid asking.
SPEAKER_00And that question is never stupid. Not once. Some of the best questions I've been asked about AI in the last year came from people with no technical background at all. Because they weren't asking how it works, they were asking what it means.
SPEAKER_02So season two is that show.
SPEAKER_00Not dumbed down. That's the distinction I want to make. We're not going shallow, we're going deep enough to give a real answer, and then stopping before the jargon starts. If the honest answer is it depends, we'll say so and explain what it depends on.
SPEAKER_02And selfishly, I think this is the harder show to make. Anyone can hide behind terminology. Explaining something clearly to someone with no background without being condescending, that's a genuinely difficult craft.
SPEAKER_00It's also the more useful show. Because the people making the biggest decisions about AI in their own lives, whether to trust it, whether their kids should use it, whether it's coming for their job, what it's actually doing with what they type into it, those people are mostly not in the industry.
SPEAKER_02So we need help. And this is the one ask of the entire season.
SPEAKER_00We need your questions. Anything you've wondered about AI and haven't had a straight answer to. How does it actually know things? Is it thinking? Why does it make things up? Is it listening to me? Should I let my kids use it? Will it take my job? Is any of this actually intelligent? What happens to what I type in?
SPEAKER_02No question is too basic. Genuinely, the more basic, the better. If you've been nodding along in meetings for two years without wanting to admit you're not sure what a model is, you are exactly the listener we're making this for.
SPEAKER_00And if you're a leader listening to this, forward the address to the people around you who've been quietly confused. Your parents. Your team. The person in your organization who's been handed an AI tool and no explanation.
SPEAKER_02The address is asksarah and james at gmail.com. That's all one word. Ask Sarah and James at gmail.com. And that's Sarah with a H. Just let us know your question, your name, and where you're from.
SPEAKER_00Send us the question you'd ask if nobody else were listening. That's the one we want.
SPEAKER_02We'll take the best ones and build episodes around them. Same two of us, same arguments about whether James is being skeptical or just lazy.
SPEAKER_00I've never once said lazy. Okay, before we go, thank you. Genuinely. Nine episodes on Data Foundations is not an obvious thing for people to listen to voluntarily, and you did. And you wrote to us. And several of you told us you'd taken something from an episode into an actual board meeting. That's the whole point of the show.
SPEAKER_02It is. And if this season did one thing, I hope it was this. The next time somebody puts an AI proposal in front of you, you now have better questions. Not more skepticism, better questions.
SPEAKER_00Fit for purpose for what? Who owns this data? Where's the capability map? Is the human in the loop enforced by code or by policy? Are we allowed to use it this way?
SPEAKER_02That's it. That's the season. Thanks for listening to AI Beyond the Hype. I'm James.
SPEAKER_00And I'm Sarah. Send us your questions. Ask Sarah and James at gmail.com. We'll see you in season two.
SPEAKER_02And remember, better AI still starts with better foundations.
SPEAKER_00Even when the foundation is a question somebody was too embarrassed to ask. Especially then.