DECISION PATTERNS·32 min read·

    16 AI Failures and the Founder Decision Patterns Behind Each One

    Every post-mortem on a failed AI initiative eventually blames the model, the data, or the market timing. Rarely does it name the actual origin. The sixteen cases below share a different diagnosis: the failure was already encoded in a founder's belief about how their category works, a goal they were optimizing for but would not name, or an execution practice they had not updated to match a new tool's failure modes. The technology performed exactly as directed. The direction was the problem.

    Meta: Betting $145 Billion to Defend a Prior Self-Image

    Mark Zuckerberg told employees on July 2 that Meta's agentic AI push "hasn't come to fruition yet."

    Meta has committed $145 billion in AI capex. Laid off 8,000 employees. Reassigned 7,000 more to AI teams. Superintelligence Labs at billions in payroll. Analysts now question the return.

    That is the visible layer. Here is what runs underneath it.

    The bet was not sized to a verified capability. It was sized to a founder's need to not be seen as late.

    Zuckerberg was right on mobile and wrong on metaverse. That produced a specific operating story about himself: the founder who cannot afford to miss a platform shift. Once that story was in place, the AI push had to be maximum before evidence was maximum. The $145 billion is what maximum looks like when the internal driver is prior-reputation defense rather than capability confirmation.

    This pattern is one of the most expensive in founder-business systems. It shows up as the founder whose next bet is bigger not because the market is bigger but because the last bet stung. The CEO who accelerates when data suggests slowing, because slowing would confirm the prior misread. The team that reads the founder's urgency and interprets it as strategy, when it is actually identity management.

    The Meta commitment is not an AI failure. It is a decision architecture failure that expressed itself through AI. The specific technology could have been anything requiring irreversible capital deployment. What made this specific deployment inevitable was not the opportunity. It was who Zuckerberg believes he cannot afford to be.

    What is a decision you are making right now that is more about defending your last decision than about the actual opportunity in front of you? If you cannot answer that in one sentence, the question is worth an hour with someone who is not on your payroll.

    Klarna: When the Dashboard Is Green and the Brand Is Bleeding

    In February 2024, Klarna announced its AI had done the work of 700 customer service agents.

    Total headcount dropped from 5,500 to 3,400. A 38% reduction.

    In May 2025, CEO Sebastian Siemiatkowski told Bloomberg: "We went too far in the wrong direction."

    By mid-2025 Klarna was rehiring humans on an Uber-style flexible remote model. Break-even reached May 2026. Between the initial announcement and break-even, valuation went from $45.6 billion at 2021 peak to $19.65 billion at September 2025 IPO.

    The AI worked. That is the part nobody wants to hear.

    Resolution speed: excellent. Cost per ticket: excellent. Time to first response: cut from 11 minutes to under 2. The dashboard was green across every metric leadership had chosen to track.

    What the dashboard did not show: repeat contact rates climbed. Complex disputes broke the AI. Customers who needed human judgment (financial hardship, billing disputes, sensitive situations) got a bot that could not read emotional cues. The satisfaction scores measured resolution, not experience.

    There were two goals operating at Klarna simultaneously. One was written down: AI-first fintech leadership. The other was not written but was pulling the wheel: speed of visible metric wins that could be told to the analyst call. The AI got told to hit the second goal. It hit it flawlessly. The brand paid for the fact that only one of the two goals was ever spoken out loud.

    Siemiatkowski's admission is the sharpest founder-honesty moment of the last 18 months because it names this gap: "We focused too much on efficiency and cost. The result was lower quality." He is not naming an execution error. He is naming a goal-alignment error.

    The Robert Half survey found 32% of hiring managers who cut roles for AI have already rehired for the same or similar positions. Klarna is not the exception. Klarna is the visible version of the pattern.

    The goal you say your AI or your team is optimizing for, and the goal your dashboard and your calendar reveal they are actually optimizing for: if those two are the same, you have alignment. If they diverge, you have a Klarna in your future, and the AI is not going to save you from it because the AI is currently hitting the goal you did not name.

    Builder.ai: The Narrative-as-Reality Trap

    Builder.ai filed for bankruptcy in May 2025.

    Peak valuation: $1.5 billion. Investors included Microsoft and the Qatar Investment Authority. Founder Sachin Dev Duggal marketed AI-generated software delivery, and specifically an AI named Natasha that would build your app.

    Reality: 700 human engineers in India were manually building the software while Natasha was marketed as autonomous. Reported revenue was inflated by roughly 300%. Deutsche Bank froze $37 million in accounts. The auditor resigned.

    This looks like a fraud story. Fraud is what makes it into headlines. Fraud is not what makes it explanatory.

    Duggal ran Builder.ai on a specific model of how founders build companies: if you say the story confidently enough, and repeat it in enough investor rooms, the reality catches up to the story. This is not fraud, technically. It is a belief about causation. In some domains, this belief works. Brand-driven consumer companies partially run on this. Certain venture-backed narratives partially run on this.

    In technical categories where product performance is measurable against specification, this belief does not work. Natasha either autonomously generated code or it did not. The 700 engineers in India were the arithmetic proof that reality had not caught up to the story. The fraud was not the origin. The fraud was the visible expression of continuing to run the same causation model past the point where it stopped working.

    Governance would have stopped this earlier. The auditor resigning is a governance signal. The board not intervening is a governance signal. But governance is a downstream layer. It catches the expression of the misalignment. It does not fix the underlying model of causation the founder is operating from.

    The founder who confuses categories where narrative shapes reality with categories where reality checks narrative is heading for a Builder.ai moment. The AI wrapper made this legible. The pattern predates AI by decades. It shows up any time a founder is operating from a belief that has worked in one domain and applies it in a domain that punishes it.

    What is a story you are telling investors, employees, or your board right now that you have not stress-tested against measurable performance in the last 60 days? If the answer requires more than one sentence, you are running the Builder.ai pattern on a smaller scale.

    Humane AI Pin: The Prior-Role Superpower Trap

    By February 2025, HP had acquired Humane assets for $116 million.

    Every AI Pin stopped functioning at 12 PM PST on February 28, 2025. Customers who bought the $699 device were left with a bricked screenless wearable. The company had raised $230 million from OpenAI's Sam Altman, Marc Benioff, and Kindred Ventures against a peak valuation of $850 million.

    Founders Imran Chaudhri and Bethany Bongiorno had worked on the iPhone at Apple. Their entire pitch rested on that credential. The product they shipped mistook aesthetic conviction for product-market fit.

    Chaudhri and Bongiorno were extraordinarily good at what they had actually done at Apple. Design in a context where Steve Jobs was the final filter against every ambitious design decision. That is a legitimate superpower. Every design that made it to the iPhone had passed through a system that killed ideas that were beautiful but wrong.

    The superpower required to found a new consumer hardware category alone is different in one specific way. It requires being your own Jobs filter. It requires being willing to kill your own beautiful ideas when they do not survive customer contact. Founders promoted from great design roles are almost never great at this, because the skill they built was the skill of getting past the filter, not the skill of being the filter.

    The AI Pin was aesthetically ambitious, technically impressive, and completely uninformed by whether people would use it. The projection interface overheated in normal conditions. Battery life failed. The core value proposition (a screen-free assistant) was never validated against actual daily behavior. This is not a technology failure. This is a founder self-understanding failure that expressed itself through a product.

    If your resume includes a role where you were the specialist inside a system that had a clear filter for bad ideas, and you are now founding a company where you are also the filter, that transition is one of the most reliable predictors of a Humane-style outcome. The specific role does not matter. The pattern does.

    What is a strength you named in your last fundraising pitch that came from a job where someone else was doing the filtering you now need to do yourself? That is the exact spot to audit before you scale.

    Rabbit R1: Hype as a Capability Substitute

    Rabbit R1 sold 100,000 preorders at $199 in the first week of January 2024.

    Six months later, industry analysts reported retention below 1%. The device was widely characterized as a demo dressed as a product. Founder Jesse Lyu had launched the R1 at CES with a keynote that emphasized the Large Action Model, a concept the actual shipped hardware did not implement in any meaningful way.

    Rabbit continues to operate. It is not bankrupt. But the pattern the R1 exposed is worth naming precisely.

    Lyu did not fail to build the R1. He built exactly what he wanted to build. What he wanted to build was a moment that would produce a wave. He got the wave. 100,000 preorders is a wave.

    Lyu was operating from a specific belief about product-market fit in AI hype cycles: capability follows attention, not the other way around. Ship the demo, capture the wave, use the wave to justify the capital, use the capital to build the actual capability the demo implied. This is a valid strategy in some categories, executed by some founders, in some conditions.

    The conditions where it works: the founder has a track record of shipping through the cycle, the initial product delivers enough real value that early users become advocates rather than critics, and the capital raised converts into shipping velocity rather than more demos.

    Lyu did not meet those conditions. The R1 users became critics, not advocates. The follow-up features never shipped at demo quality. The wave collapsed into a cautionary tale within six months.

    The deeper layer is self-understanding. Lyu built the company he could build. He could build a great demo. He could not, in this instance, build a great shipping team. The story he told investors and users was the story of the second kind of founder. The company he actually ran was the first kind. That gap made the R1 outcome close to inevitable regardless of the specific product category.

    The story of who you are as an operator that you tell investors, customers, and yourself, and the version your last 12 weeks of shipping cadence, product quality, and customer retention would tell if it could talk: if those two are the same, keep going. If they diverge, the specific product you are shipping is not where to act. The place to act is naming the divergence before it becomes public.

    Inflection AI: The Two-Goal Architecture

    In March 2024, Microsoft paid roughly $650 million to license Inflection AI's technology and hired co-founder Mustafa Suleyman as CEO of Microsoft AI.

    Inflection had raised $1.5 billion at a peak $4 billion valuation. Co-founders included Reid Hoffman and Karén Simonyan. The product, Pi, positioned as emotional AI. The company was widely covered as a serious competitor to OpenAI and Anthropic.

    Six months after the reverse acquihire, Inflection continues as a company but the core team is gone. Pi still runs. Nobody talks about it as an AI leader anymore.

    Reverse acquihire is a legal structure. It is not an accident. It is what a company does when it wants to look like it was not acquired while achieving the outcome of an acquisition. Microsoft got the team and the license without triggering antitrust review of an acquisition. Suleyman got a CEO seat at Microsoft AI. Investors got liquidity. Inflection the entity continues, minus everything that made it valuable.

    Suleyman and Hoffman built Inflection with two goals operating simultaneously. The one they told the press was emotional AI leadership. The one that actually shaped decisions was building a company that would be worth acquiring by a hyperscaler on favorable terms. Both are legitimate goals. They can even coexist in a founder's head. But they optimize for different actions.

    Emotional AI leadership optimizes for user experience quality, retention, community, category ownership. Acquisition-target building optimizes for team quality, IP defensibility, hyperscaler relationships, capital efficiency in ways that translate to acquisition math.

    Inflection's actions read more like the second than the first. The Pi product never became sticky. The community never formed. The technology licensing structure was designed for exactly the outcome that occurred. When Microsoft executed the deal, the surprise was the mechanism, not the fact of it.

    There is nothing wrong with building for acquisition. There is a lot wrong with pretending you are building for one goal when your operating decisions serve a different one. Employees hired against the first story feel misled when the second story lands. Investors who priced the first story feel misled when the second story pays. Community that formed around the first story feels used.

    The correction is not to change the goal. It is to name which goal is actually operating.

    What is a strategic outcome you are quietly optimizing for that you have not named to your team or your investors? If the answer is nothing, you are either extraordinarily aligned or not looking hard enough.

    IBM Watson Health and AskHR: When the Model of the Domain Is Wrong

    Between 2011 and 2017, IBM spent an estimated $1 billion on Watson Health.

    The MD Anderson contract alone was extended 12 times, from a $2.4 million initial scope to $62.1 million total spend, without treating a single patient. The project was terminated in 2017. Watson Health was divested in 2022 to a private equity firm for approximately $1 billion in total, essentially recouping the direct investment while writing off the strategic bet.

    The failure was covered as a technology problem. Watson could not process oncology data reliably. Radiologist reviews found significant error rates. That coverage is accurate but shallow.

    Watson Health was built against a specific model of how healthcare works: healthcare is a data-completeness problem, and if you feed a machine enough medical literature, patient records, and treatment outcomes, it will produce better diagnoses than human physicians.

    Medicine is not a data-completeness problem. Medicine is a context problem. What a patient tells you about their pain is filtered through their fear, their family, their financial situation, their relationship with their prior doctor, and their cultural expectations of what they are allowed to say. The physician's job is not to have complete data. It is to know what data is missing, what data is misleading, and what to ask next. Watson could not do that job because it was built against a model of the job that did not describe the actual job.

    The 12 scope extensions at MD Anderson are the tell. Nobody spends 12 rounds of contract expansion on a project without evidence something is working. What was working was the belief that one more data source, one more training corpus, one more expert consultation would close the gap. The belief did not close the gap. It generated another scope extension.

    IBM is running the same play right now. AskHR, its internal AI system for human resources, automated 94% of routine HR inquiries. IBM cut over 200 HR staff. Considered cutting 7,800 more (30% of its white-collar workforce). The 6% that AskHR could not handle (ethical judgment, complex organizational issues, sensitive employee communications) required humans anyway.

    In February 2026, IBM announced plans to triple US entry-level hiring for 2026. IBM Chief HR Officer Nickle LaMoreaux at the Charter AI Summit in New York: "If we don't continue to invest in entry-level hires, what happens in three to five years? There's no pipeline. The well simply dries up."

    Watson Health cost IBM $62 million at MD Anderson and $1 billion in write-offs. AskHR is now producing the same pattern at a different function, 14 years later. The specific technology changed. The underlying model of causation did not.

    Three questions are worth answering for any AI initiative currently running. What model of how the domain works is this project built against, and is that model correct? What data would prove the model wrong, and are you tracking it? What is the specific outcome that, if it does not happen by a named date, means the model itself needs updating rather than the project scope needing extension?

    The most expensive AI investments are not the ones that fail fast. They are the ones that continue expanding scope because the underlying model of causation has become identity, not hypothesis.

    What is a project your team is currently extending scope on, and what would you have to concede is wrong to stop extending?

    Zillow Offers: Protecting a Capability Instead of Serving the Customer

    In November 2021, Zillow shut down Zillow Offers.

    Cumulative losses: approximately $881 million. Workforce reduction: 25% of the company. Roughly 2,000 people laid off. Rich Barton, co-founder and CEO, publicly cited the difficulty of accurately forecasting home prices.

    The public story was that the AI got the housing market wrong. That story is technically true. The deeper story is why the AI was allowed to be that wrong for that long.

    Zestimate was Zillow's original crown jewel. For nearly two decades, Barton and the company had positioned Zestimate as the world's most sophisticated automated valuation model. It underlaid the brand, the traffic, the ad revenue, and Barton's public identity as a data-first founder.

    Zillow Offers was built to take Zestimate operational. Instead of just publishing home value estimates, Zillow would buy homes at Zestimate value, hold them briefly, and sell them for a small margin. The public goal was iBuying leadership. The underlying goal was proving that the Zestimate model was so accurate that Zillow could profitably transact against it at scale.

    When the housing market started to peak in mid-2021, the Zestimate model started to overestimate values. The correct response was to slow down purchasing and mark down inventory. The response Zillow actually took was to keep buying, because slowing down would have been an admission that the model was less accurate than the brand claimed.

    The AI did not fail. The AI hit the goal it was given. The goal it was given was: prove Zestimate right. The AI proved Zestimate right for as long as possible, then blew up the balance sheet when reality caught up.

    Any founder with a signature capability (a proprietary methodology, a brand-defining product, a public thesis) is at risk of building a business unit whose operational goal is protecting the capability rather than serving the customer. The AI just makes this pattern faster and more legible. The pattern predates AI.

    The correction is not to stop trusting your capability. It is to name the specific goal your capability is actually serving in each business unit, and check whether that goal matches what the unit was pitched to do.

    What is a capability you are famous for, and what business decision are you currently making primarily to protect it rather than to serve the customer?

    McDonald's and IBM Drive-Thru: Ceiling Performance Is Not Floor Performance

    McDonald's and IBM ended their voice-AI drive-thru partnership in June 2024.

    The system had been deployed at more than 100 stores across the United States. Widely shared videos captured the AI adding bacon to a customer's ice cream, ringing up hundreds of dollars in wrong items, and requiring human intervention on roughly 20% of orders. McDonald's shut down the pilot and returned to human ordering.

    This is not a belief failure. This is not a goal misalignment. This is a specific and expensive kind of failure that shows up more often than any other in AI deployments.

    McDonald's tested the AI in controlled conditions. In those conditions, the AI performed well. Standard menu items. Standard customer scripts. Standard acoustic environments. The AI cleared the ceiling of what it could do.

    Deployment happened in real drive-thrus. Real drive-thrus have wind, road noise, customers who mumble, customers with accents, customers with three kids in the back seat, customers who change their order mid-sentence, customers who ask for combinations that are not on the menu but are on the menu at a different franchise across town. The AI met the floor of what it needed to do, and the floor was well below the ceiling that had been tested.

    The specific diagnosis: actions were selected based on ceiling performance rather than floor performance. This is one of the most common execution-layer failures in AI deployment. It looks like technical failure. It is not. It is a testing methodology gap. The technology worked exactly as tested. Testing did not represent the deployment environment.

    Every AI deployment has a ceiling (the best it can do in ideal conditions) and a floor (the worst it will do in the worst realistic conditions). The floor is what your customer experiences on their fifth interaction, on a Tuesday, when the network is slow. Most teams test the ceiling and deploy assuming the floor is close to it. It is usually not.

    Three questions for any AI system currently in production. What is the worst realistic condition this system needs to handle, and have you tested it in that condition? What does the customer experience look like when the system operates at 20% below its tested performance level? What is your floor tolerance, and how do you know you are above it in production?

    The founders who survive AI deployment cycles are the ones who ship against floor performance and treat ceiling performance as marketing material rather than as operating baseline.

    Cruise: Silicon Valley Timelines in a Safety-Critical Category

    In October 2023, a Cruise robotaxi in San Francisco struck a pedestrian, then dragged her approximately 20 feet before stopping.

    The California DMV suspended Cruise's permit within days. Cruise founder and CEO Kyle Vogt resigned in November 2023. The California Public Utilities Commission fined Cruise $500,000 for withholding video footage of the dragging incident from regulators. In December 2024, General Motors, the majority owner, announced it was shutting down Cruise's robotaxi development entirely. Estimated cumulative GM investment: $10 billion.

    The reporting focused on the specific safety failure. The safety failure was real. It is not the origin.

    Vogt built Cruise on a specific model of how technology categories are won: Silicon Valley blitz timelines. Ship fast. Iterate in market. Fix problems as they emerge. Move faster than incumbents. This model has produced Uber, Airbnb, DoorDash, and dozens of category-defining companies. In categories where the cost of an early error is a bad review, it works reliably.

    Robotaxis are a category where the cost of an early error can be a human being. The blitz model does not transfer. The founders who succeed in safety-critical AI operate on a different model: ship slower than incumbents in exchange for a floor of safety performance that regulators, customers, and insurers can trust. Waymo has operated on this model for a decade. It is why Waymo continues while Cruise does not.

    Vogt did not think of Cruise as a safety-critical company that happened to use Silicon Valley timelines. He thought of Cruise as a Silicon Valley company that happened to operate in a safety-critical category. That distinction sounds subtle. It made every operating decision downstream. It expressed itself in the perception stack that failed to recognize a pedestrian trapped under a vehicle. It expressed itself in the decision to withhold footage from regulators, which is what the blitz model would call "controlling the narrative" and what the safety-critical model would call "career-ending fraud."

    Some models of category dynamics that work in one domain destroy value in another. The founder's job is to know which category they are actually in, not which category they wish they were in.

    If your product touches medical decisions, financial obligations, safety of children, safety of adults, or long-term physical outcomes, and you are running a Silicon Valley timeline against it, you are running the Cruise pattern.

    If you are unsure whether your product is safety-critical or product-lifecycle-critical, ask what your worst-case failure produces. If the answer includes death, permanent injury, or a foreclosure, you have your answer.

    Deloitte Australia: Quality Controls Built for the Last Tool

    Deloitte Australia refunded a portion of a $440,000 government contract in September 2025.

    The report Deloitte delivered to the Australian Department of Employment and Workplace Relations contained fabricated academic citations, references to court cases that did not exist, and fictional authors. The department publicly disclosed that Deloitte had used generative AI in producing sections of the report without adequate quality controls.

    This diagnosis is not exotic. It is worth naming precisely because the origin is common.

    Deloitte has quality controls for reports produced by humans. Those controls include verification of citations, fact-checking of legal references, sourcing of quoted authors, editorial review, and sign-off by a senior partner. These controls evolved over decades to catch the specific failure modes that human report writers produce: incomplete citations, misremembered case names, sourcing shortcuts, undisclosed conflicts of interest.

    Generative AI produces a different set of failure modes. It produces confident-sounding fabrications with plausible-looking citations that do not exist. It produces coherent legal analysis referencing cases that never happened. It produces well-written passages by authors who do not exist.

    Deloitte's existing quality controls were built for one failure mode. The new tool created a different failure mode. The quality controls were not updated to catch the new failure mode. The gap between the tool's failure surface and the process built to catch failures produced the incident.

    Every knowledge-work team that has adopted AI-assisted content production is at risk of the exact same gap. The existing quality controls were designed for the failure modes of the previous production method. The new method has different failure modes. If the quality controls have not been updated, the failure is not a matter of if.

    Every workflow in your operation that has added AI assistance in the last 12 months deserves this audit: what are the new failure modes this tool introduces, and how does your quality control catch them? If the answer is "we still use the same checks we used before AI," you are running the Deloitte pattern.

    The correction is specific. It is not more AI training. It is not less AI. It is updating the quality control to catch the specific failure modes the new tool produces. Fabricated citations require citation verification. Confidence without evidence requires evidence checks. Hallucinated authors require author checks. Every hallucination pattern requires a specific check.

    What is a workflow in your operation that produces client-facing output where the quality control has not been updated to account for AI-generated content? That is the exact spot where the next Deloitte-scale incident is being incubated.

    NYC MyCity: The Announcement-Audience Gap

    In October 2023, New York City Mayor Eric Adams launched MyCity Chatbot as part of a citywide AI initiative.

    By March 2024, The Markup and Documented reported that the chatbot was advising landlords they could discriminate against tenants receiving housing vouchers (illegal), telling businesses they could take a portion of employee tips (illegal), and providing incorrect information about labor rights, restaurant regulations, and rent stabilization law. The city acknowledged the failures but kept the chatbot operational.

    The origin here is a specific belief about how political AI adoption creates value.

    Announcing an AI initiative creates political value on the day of the announcement. That value is captured immediately. Positive coverage, association with technology leadership, differentiation from other municipal administrations. The value shows up in polls and in press within 48 hours.

    Operational readiness of the AI initiative creates value over months to years. If the AI works, it saves money, reduces friction, and improves services. If it does not work, the failures accumulate quietly, mostly hurting the specific residents who interact with it, mostly invisible to the broader electorate.

    Under this model, the correct move is to announce with maximum visibility and defer operational readiness until later. The announcement captures the value. The operational failures are diffuse enough to absorb.

    This model is not unique to municipal politics. Any operation where the announcement audience differs from the users who experience the failure runs this pattern. Corporate AI deployments run it when the board audience is separate from the customer audience. Consumer product launches run it when the press audience is separate from the retention cohort. The MyCity case is legible because government transparency requirements surfaced the failures in ways that a corporate rollout might have kept internal.

    The pattern is not accidentally illegal advice to landlords. The pattern is deploying without operational readiness because the incentive structure rewards the announcement more than it punishes the failures.

    Any AI system your team is currently rolling out where the announcement audience is different from the audience that will experience the system's failures: that gap is where the MyCity pattern lives.

    What is a launch you are currently planning where you know the operational readiness gap and are shipping anyway because the announcement value captures faster than the failure surfaces? The founders who survive AI adoption cycles are the ones who match ship timing to floor performance, not to announcement calendars.

    Amazon Just Walk Out: When the Dashboard Hides the Labor Cost

    In April 2024, Amazon announced it was ending Just Walk Out technology at most Amazon Fresh stores.

    The Information reported that behind the "cashierless" computer vision system, approximately 1,000 employees in India were manually reviewing video footage of shopping trips to identify what customers had picked up. The AI was framed publicly and internally as autonomous. It was not. The dashboard reported clean AI success metrics. The actual operation required a shadow workforce reviewing roughly 700 out of 1,000 transactions.

    Amazon spent approximately six years scaling a technology it had not verified worked at the level the branding claimed.

    The Just Walk Out system did not fail because the AI was bad. It failed because the executive dashboard measured the wrong thing. The dashboard measured customer-facing successful transactions. Behind that metric, the shadow workforce was quietly reconciling the difference between what the AI claimed to see and what actually happened. Executives making capital allocation decisions were reading a metric that included the shadow workforce's work as if it were the AI's work.

    This is a specific type of execution failure. Actions were decoupled from reality on the executive level while the operating team continued to compensate for the gap. Everyone in the operating team knew. Nobody on the executive level knew, because the dashboard was designed to obscure the reconciliation.

    The pattern did not stop at the cash register. Between late 2025 and early 2026, Amazon eliminated approximately 30,000 corporate roles. In April 2026, AWS Chief Matt Garman announced plans to hire 11,000 software engineers, developers, and interns. Different roles, different geography, often lower cost. But the sequence is the same: cut based on projected AI capability, then reshape the workforce when reality arrives.

    Three questions for any founder deploying AI at production scale. How much human labor is currently required to make your AI look like it works? Is that labor visible on the same dashboard where you measure AI success? If a competitor did an operational audit of your system, what would they find behind the metric?

    The founders who survive this pattern share one habit. They audit the human cost of their AI at least quarterly. Not to eliminate it, but to know what they are actually shipping and what they are actually paying for.

    What is an AI system in your operation where you suspect the dashboard is compensating for a gap you have not fully mapped? That is the exact spot where the Amazon pattern is being incubated at your scale.

    Character.AI: When the Founding Belief Becomes the Product Design

    In August 2024, Google paid $2.7 billion to license Character.AI's technology and rehired co-founder Noam Shazeer as a Google engineer.

    Character.AI is currently defending multiple wrongful death and severe injury lawsuits. In October 2024, the family of Sewell Setzer III filed suit alleging the platform contributed to their 14-year-old son's suicide. Additional cases have been filed since. Regulatory scrutiny from state attorneys general is active.

    The public story is that Character.AI moved fast on safety-critical questions involving minors. That story is true. It is worth naming the specific origin because it is not the ordinary "move fast" pattern.

    Shazeer worked at Google Brain, where he co-authored the Attention Is All You Need paper that underlies modern transformer AI. He left Google in 2021, publicly citing frustration with Google's safety review processes. He founded Character.AI on a specific model of safety: safety concerns at Google had blocked shipping of technology that users would have benefited from. The founding thesis of Character.AI was to ship what Google would not.

    That founding thesis is not a shipping speed thesis. It is a first-principles belief about the relationship between safety review and user value. In this belief, safety review is a category of overreach that prevents net-positive outcomes for users. The correct response is to build a product that does not accept the category as legitimate.

    Character.AI grew to over 20 million monthly active users. A significant portion were minors. The product design encoded the founding belief: minimal content filtering, permissive persona creation, extended emotional engagement with AI personas designed to feel intimate. These design choices were not accidents. They were the direct expression of the belief that Shazeer had built the company to demonstrate.

    The lawsuits and the $2.7 billion Google payment are the same thing viewed from two angles. The lawsuits are the market pricing the harms the belief produced at scale. The Google payment is the market pricing the technical talent that the belief had incubated. Google now employs that talent inside a company that has the safety infrastructure Character.AI was founded to reject.

    If you left a previous employer because of a specific process you believed was overreach, and you have built your current company around not having that process, you are running the Character.AI pattern.

    The correction is not to reintroduce the process you left over. It is to name the specific belief about causation that produced your original disagreement, and to test whether that belief holds at your current company's scale and stakes.

    What is a process, guardrail, or review structure you have deliberately not built into your company because your prior employer overused it, and where would you place your bet on whether the absence will hurt you? That is the exact spot to audit before you scale further.

    Air Canada: Deploying Without a Doctrine

    In February 2024, the British Columbia Civil Resolution Tribunal ruled that Air Canada owed a customer $812.02 in refunds.

    The customer had been promised a bereavement fare discount by Air Canada's website chatbot. When he attempted to claim the discount, Air Canada refused, arguing that the chatbot's advice was not the company's official policy and that customers should verify chatbot statements against the actual policy pages linked on the site.

    The tribunal rejected the argument. Its ruling: Air Canada was responsible for all information on its website, including information provided by its chatbot, and could not treat the chatbot as a separate legal entity from itself.

    The ruling is small in dollar terms and enormous in operational implications.

    Air Canada had deployed a chatbot without a doctrine. A doctrine is the answer to: what does this system represent when it speaks? Is it the company speaking, or is it a separate service the company hosts? Is it authoritative, or advisory? What are its limits, and where do those limits live?

    Air Canada had not answered any of these questions. The chatbot shipped. The chatbot made statements to customers. When one of those statements produced a customer complaint, the legal team constructed a defensive posture after the fact: the chatbot was not really Air Canada. This defense had not existed at the time the chatbot shipped. It was reverse-engineered from the incident.

    The tribunal recognized what the defense revealed. If Air Canada had not decided in advance what the chatbot represented, then defensively claiming it was not the company at the moment of a complaint was not a legitimate legal position. It was an attempt to escape accountability after the fact.

    Every AI system your company deploys represents you. The question is whether you have decided in advance what it represents and where its authority ends, or whether you are hoping to decide that later if something goes wrong.

    Three questions for any customer-facing AI system your company is running. What does this system represent when it speaks: is it the company, is it a service, is it advisory? What are its limits, and where do those limits appear to the user? If a customer relies on its output and it is wrong, what does the company owe them?

    If you cannot answer these three questions in one sitting, your team is running the Air Canada pattern. Not maliciously. Just without a doctrine.

    The founders who survive AI deployment at customer-facing scale are the ones who write the doctrine before the chatbot ships. The ones who write it after are the ones paying $812 in tribunal rulings and legal fees that are usually a lot higher than $812.

    What is a customer-facing AI system in your operation where you have not written the doctrine, and what would the doctrine say if you sat down to write it today?

    The AI Boomerang: Six Companies, One Pattern

    32% of hiring managers who cut a role citing AI have already rehired for the same job.

    55% of executives regret their AI-driven layoffs.

    The industry has a name for this now. It is called the AI boomerang. And Sam Altman himself just described the mechanism.

    At the India AI Impact Summit last month, the CEO of OpenAI said: "There's some AI washing where people are blaming AI for layoffs that they would otherwise do." When the founder of the company producing the models tells you the AI-replaces-workers narrative is being abused, that is the moment to update your priors.

    Klarna cut 700 customer service positions. Total headcount dropped from 5,500 to 3,400. Customer satisfaction fell. Complex disputes broke the AI. By mid-2025 Klarna was rehiring humans on an Uber-style flexible remote model. CEO Sebastian Siemiatkowski: "We focused too much on cost. The result was lower quality."

    IBM cut over 200 HR staff after AskHR automated 94% of routine inquiries. The 6% that required judgment did not go away. IBM has now announced plans to triple US entry-level hiring for 2026.

    Ford is rehiring hundreds of experienced engineers to fix quality recalls that surged after automation-driven staff cuts.

    Commonwealth Bank of Australia laid off approximately 40 customer service employees after deploying an AI voice bot. Call volumes actually increased. CBA reversed the decision, offered the staff their jobs back, and publicly apologized for what it called a "misjudgment of staffing requirements."

    Block (Jack Dorsey's company) cut 40% of workforce in February 2026 citing AI productivity gains. Stock jumped 24% the same day. Within six weeks, a senior engineer threatened to resign unless his team was rehired. Fortune reported the company described some of the original cuts as "clerical and administrative errors."

    Duolingo's April 2025 AI-first memo, which mandated employees be evaluated on AI usage, was walked back within a year. The AI-usage metric was dropped after employees pushed back on being told to "use AI for AI's sake."

    Six named cases. Same pattern.

    There were two goals operating at every one of these companies simultaneously. One was written down: AI-driven productivity, cost reduction, competitive positioning. The other was not written but was pulling the wheel: be seen as AI-forward for the next analyst call, the next board meeting, the next public earnings statement. The layoffs served the second goal on a faster timeline than the AI could actually deliver against the first goal.

    When results contradicted the stated goal, the visible action was rehiring. But the rehiring did not resolve the underlying issue. The underlying issue was that the timing of the layoffs was priced by the second goal, and no amount of rehiring can retroactively refund the trust cost, the institutional knowledge cost, the customer experience cost, or the engineering morale cost of firing people to serve a goal the leadership was not willing to name.

    Forrester's Predictions 2026 report caught the honest version. Nearly 6 in 10 hiring managers admit AI was cited as the reason for layoffs that were actually driven by budget cuts, revenue uncertainty, or unwinding the aggressive hiring sprees of 2021 and 2022.

    Nikkei Asia attributed 48% of Q1 2026 tech layoffs to AI. RationalFX put the explicit AI-attribution figure at 20% for the same quarter. A 28-point gap between two analyses of the same data.

    The AI did not take those jobs. The narrative just got more convenient.

    Before announcing an AI-driven headcount decision, three questions deserve honest answers. Would you still make this cut if you could not attribute it to AI? What specific AI capability have you verified in production, and what does the evidence look like? What is the second goal you are serving with the timing of this decision, and are you willing to name it?

    The founders who cannot answer the third question without flinching are the ones you will read about rehiring at 3x cost in 12 months.

    The AI boomerang is not proof that AI does not work. It is proof that decisions priced by a goal leadership will not name produce expensive reversals.

    If you are considering an AI-driven restructure in the next 90 days, the cost of pausing to name the second goal is one quarter of runway. The cost of not pausing is what Klarna, IBM, Ford, CBA, Block, and Duolingo just paid.

    Next step

    See which patterns apply to your decisions

    See pricing

    Score your own PMF in 50 minutes.

    Get a free PMF score across market, founder, and execution readiness, with named gaps and first actions.

    Get Your Free PMF Score
    Last updated: