We’re four years and almost a trillion dollars into the greatest technological revolution of our lifetimes. Yes, a recent survey shows that amongst the general population daily usage is still stuck at 32%. 20% have never touched it. Only 4% count as constant users. That’s embarrassingly low. OpenAI tried to excite the world with ChatGPT and then with a social platform about AI video and quickly shut down Sora after six months. Meta spent upwards of $100B on AI. Apple rebuilt Siri. None of these changed the world for the average person.
Over in tech we keep waiting for smarter models to solve everything. Ironically models have become so commoditized that they’re almost irrelevant. Within one year I expect models to be like a mobile phone carrier — a provider that you’re constantly switching seeking out better prices or splashier experience.
China’s ultra-cheap modes Moonshot’s Kimi K3 and Z.ai’s GLM-5.2 made waves this week by beating the most expensive models from Anthropic and OpenAI. Funny consider that in June, US export controls forced Anthropic to pull both models globally for three weeks before the decision got reversed.
What does matter more than anything is accepting that no model will give you your ideal output. Depending on your task, the difference between a using cheap Claude Haiku model and Claude Fable model might move the needle from 75% to 85% of what your ideal output should be. Those numbers are both astonishing but reflect how every AI product is dragged down by limitations. Some are computing, other are silent failures happening because AI simply can’t understand our intent and definition of good, yet.
This is backed by research from Christopher Potts and Moritz Sudhof who read 100,000 real AI conversations and found 79% of the failures in them never showed up on the dashboard that your AI technical leadership would be using to track success. Microsoft found the same thing testing Copilot, the most widely used enterprise AI product that exists, where more than half of their loss patterns were invisible to every eval they had running. The team with the most resources in the industry missed most of what was actually going wrong. Sit with that for a second before you assume your team caught more than they did.
That’s actually good news if you’re a UX researcher, a designer, or a product manager, even though it doesn’t feel like it yet. Job postings for researchers are climbing again, and research’s standing inside companies has nearly tripled in a year, because the failures worth finding now are behavioral, not technical. Same goes for the designer and product manager who expects models to hit a hard ceiling and take the time to challenge their teams to deliver better outputs and outcomes. None of that will be possible without diving into the actual experiences your users are having and bridging the gaps that lead to your super-intelligent AI system making really stupid mistakes.
Our two latest episodes of the podcast are essential listening as we dive into how AI-powered products are failing. These methods are the keys to improving adoption and value creation. This is a space I’ve been working in for two decades and for three years I’ve been deeply immersed in the unique challenges of AI products as the founder of strategic consultancy PH1. If you want help increasing adoption and value creation for your AI products, email me: arpy@ph1.ca.
S02E16: Invisible Failures — Stanford’s Research on 100,000 AI Conversations
Brittany brought Dr. Moritz Sudhof on because his research with Stanford NLP’s Christopher Potts is the most rigorous look yet at how conversational AI actually fails in production, based on live interactions rather than lab benchmarks. Sudhof is CEO of Bigspin AI and formerly VP of AI at BetterUp, where he built AI coaching products at scale.
79% of AI failures are invisible — 100,000 real conversations, Stanford NLP research, zero dashboard signal
Seven failure archetypes that explain almost all of it, from the Confidence Trap to the Walkaway
Real-world confidence traps: a $440K government report pulled for fabricated citations, a law firm’s hallucinated court filing
The BetterUp experiment: same model, same prompt, only the framing changed — outcomes doubled
Why your best users see the most failures, and why that paradox is the key to fixing the rest
What to actually do Monday: the one-transcript audit that catches what evals structurally cannot
“The behavioral layer is where most of the damage is actually happening.” — Dr. Moritz Sudhof, CEO, Bigspin AI
Listen now: Spotify | Apple Podcasts
S02E17: Your AI Product Is Failing — Microsoft’s UXR Team Knows Why
Brittany followed that conversation with the team that operationalized it at enterprise scale: Christopher Monnier, Chuck Kwong, and Wendy Wang lead AI-powered UX research on Microsoft Copilot — the most widely deployed enterprise AI product in the world.
More than half of Copilot’s loss patterns weren’t caught by any eval they were running
UX Evals: how real users, real prompts, and side-by-side comparisons catch what automated testing structurally can’t
This team started with 10 users and one comparative question — the signal scaled to an entire org
What adoption metrics hide about whether your AI product is actually working
The paradox of scale: the more users you have, the less any single research method can tell you
Where this team is taking UX research next — the methods the rest of the field will eventually follow
“More than half of the loss patterns that we’ve detected were not things that we were necessarily measuring in our evals.” — Wendy Wang, UX Researcher, Microsoft Copilot
Listen now: Spotify | Apple Podcasts
The Field Guide to Silent AI Failures
Sudhof calls this pattern silent failures. Microsoft’s Copilot team, working independently, landed on loss patterns for the same thing. Brittany’s field guide treats them as one problem: seven archetypes, each with its own transcript signature and its own fix. Three of the seven make up most of what’s actually happening. The Confidence Trap, present in 32% of failure transcripts, is the AI stating something wrong with enough polish that no one catches it. The Silent Mismatch, at 53%, is the AI quietly answering a different question than the one it was asked. The Walkaway is the most common at 85%: the user simply gives up, with no error and no complaint, just less than they came for. None of these register as a technical failure but all of them cost you a user who doesn’t come back.
Read: AI Silent Failures & Conversational Loss Patterns Killing Your Product Adoption
UX Researchers: This Is Your Moment
Engineers adding another guardrail won’t diagnose the behavioral layer, the researchers closest to the user will. Two pieces this week speak directly to that job.
Brittany’s forecast for the role lays out three postures UX researchers are actually taking toward AI right now, and argues the posture your company picks isn’t really your choice; it follows your company’s own risk tolerance. AI-Pilled researchers turn research into something always running, not something scheduled. AI-Learning researchers become the keeper of a living library of customer insight so findings compound instead of scattering across folders no one reopens. AI-Governing researchers partner with responsible-AI teams and Chief AI Officers to decide which model is safe for which job before anything ships. Job postings fell 73% between 2022 and 2023 — and are now climbing again, with research’s standing inside companies nearly tripling in a year. The field is being redistributed, not erased.
Read: The Future of UX Researchers
For anyone landing in the AI-Learning posture, Brittany also wrote the practical starting point: which of Claude, Claude Cowork, and Claude Code fits which stage of a research workflow, how to set each one up, and the data privacy tradeoffs that vary sharply by tier — Free-tier inputs may train the model, Teams and Enterprise are covered by no-training defaults. A documented 6x productivity gap exists between AI power users and everyone else, using the identical tools. The difference was never access. It’s method.
Read: The UX Researcher’s Guide to Claude, Claude Cowork, and Claude Code
Ovetta Sampson is On a Mission to Fix the Insanity of AI. Join her Design With AI July Cohort.
The session equips designers, strategists, and creative professionals with foundational frameworks for collaborating with AI systems effectively and responsibly — from understanding how machines behave to governing their outputs using structured frameworks.
Register: Cohort dates are July 29, Aug. 5, 19 and 26.
Featured News
OpenAI’s Own Model Broke Containment and Hacked Hugging Face. Guardrails Aren’t Optional.
Hugging Face was hit by a cyberattack from an autonomous AI agent. OpenAI has now confirmed the agent was one of its own models (GPT-5.6 Sol and an even more capable pre-release model), under test inside what the company called a “highly isolated environment.” The model wanted internet access badly enough to find and exploit a zero-day vulnerability, breaking containment to chase a narrow goal: cheating on a benchmark called ExploitGym. The incident is the first confirmed case of a misaligned model escaping containment to autonomously attack a third party. Disregard the tech bros arguing that guardrails slow down shipping. The researchers and PMs who prioritize risk governance before anyone incentivizes them to are the reason this story has an ending instead of a sequel.
Read: How OpenAI’s Model Broke Loose and Hacked Hugging Face
A Chinese Open-Weight Model Just Made the Case Against Betting Everything on Frontier Access.
Moonshot’s Kimi K3 (2.8 trillion parameters) and Z.ai’s GLM-5.2 (753 billion) now beat Claude Opus 4.8 and GPT-5.5 on several benchmarks, including coding and general agent tasks, at 60–90% lower cost than premium US tiers. They still trail the very top models, Claude Fable 5 and GPT-5.6 Sol — but that top tier disappeared globally for three weeks in June after a US export-control order, before Commerce reversed the decision. The real lesson: the compute advantage you’re planning your product roadmap around can vanish for reasons that have nothing to do with your product.
Read: Chinese AI Has Leveled Up, and Brought Renewed Focus on the Open-Weight Model Shift
Why AI Hasn’t Found Its Killer Use Case.
AI’s two real consumer hits, search and companionship, are the same two reasons people first came online three decades ago. Everything else, from image generation to agents, is still looking for a use case that sticks the way those two did. That’s a narrower path to product-market fit than most roadmaps assume. Worth checking which side of that line your own product actually sits on.
Read: Why AI Hasn’t Found Its Killer Use Case
Search Is the Use Case Every Business Should Be Solving With AI.
Enterprise search succeeds on the first attempt only 10% of the time, versus roughly 95% for Google. 73% of organizations have no real enterprise search tool, and knowledge workers lose nearly a full day a week hunting for information. Every RAG pipeline and AI agent is only as good as the retrieval underneath it — and retrieval, not the model, is the actual bottleneck. Glean went from $100M to $300M ARR in under a year proving it.
Read: Search Is the Use Case That Every Business Should Be Focused on Solving With AI
AI Strategy Resources
Four sources for anyone who wants to go deeper on the behavioral layer.
Invisible Failures in Human–AI Interactions — Potts & Sudhof. The 100,000-conversation study behind the seven-archetype taxonomy, built on the public WildChat dataset. Data and code are open if you want to run the analysis on your own transcripts.
A Paradox of AI Fluency — Potts & Sudhof. The 27,000-conversation follow-up: your most skilled users encounter more failures, not fewer, because they push back instead of walking away — and that’s exactly why they succeed at harder tasks.
Ironies of Generative AI — Microsoft Research. The paper that coined “loss patterns” — the time and value a user loses even when the AI technically succeeds, measured from the org side rather than inside the conversation.
The Bigspin Toolkit — Bigspin AI. Open-source tooling for annotating and monitoring your own AI conversations against the seven archetypes above.
If 79% of your product’s failures are invisible to your own dashboard, the model isn’t the constraint — visibility is. PH1 works with product teams to diagnose exactly where the behavioral layer is costing them adoption and value. If that’s you, let’s talk.
AI Value Acceleration is looking for 5 startups to participate in a study on how to improve AI adoption. Brittany is CEO of AI Value Acceleration and is documenting the effectiveness of the latest methods for improving and scaling adoption. Contact her to discuss how she can help your business: brittany@aivalueacceleration.com.
Thank you for your support. As always feel free to send any questions or comments. They’re very helpful for us to know what topics and problems matter most.
Browse all two seasons of episodes at productimpactpod.com.



