arrow-up

Can AI Replace Humans? We Tested It


Can AI replace human experts? We tested it, and the errors follow a pattern

In October 2025, Deloitte Australia refunded part of a AU$440,000 fee to the Australian government over a report containing citations to academic papers that do not exist and a fabricated quote from a federal court judgment, produced with help from GPT-4o. Everyone in the AI debate uses stories like this to argue that people still matter. Almost nobody explains the part that actually matters for a hiring decision. These failures follow a pattern, and the pattern tells you exactly what you are paying an expert for.

So can AI replace humans? We are an SEO agency that uses AI daily, so we ran the test ourselves. We gave an AI model a one-line prompt asking for a complete, common deliverable in our field, then audited the output against the platform rules we deploy against every week. The result was 10 blocks of professional-looking markup containing 14 validation, policy, or structural failures and 36 fabricated data points, including a board-certified doctor who does not exist. Every failure was predictable in advance. This article shows the test, the pattern behind it, and the research confirming that pattern holds far beyond SEO.

Key takeaways

  • AI errors cluster in two predictable places. Where the world changed after the model’s training, and where your situation differs from the average case.
  • Our one-prompt test returned 14 validation, policy, or structural failures and 36 invented facts, including a fabricated 4.8-star rating, a named patient testimonial, and a fully credentialed Medical Director who does not exist.
  • In a Harvard study of 640 entrepreneurs, the identical AI advisor raised strong performers’ profits about 15% and cut weak performers’ profits about 8%.
  • Consultants using GPT-4 beyond its ability were 19 percentage points less likely to reach the correct answer. Inside its ability, the same consultants finished 12% more work 25% faster.
  • AI collapsed the cost of producing work. It did nothing to the cost of verifying it. That gap is what experts now get paid for.

Can AI replace humans in skilled work?

No. AI produces expert-sounding output, but its errors concentrate where expertise lives, in what changed recently and in what makes your case different from the typical one. Research from Harvard, MIT, and BCG shows AI amplifies the skill of its user. Skilled users get leverage. Unskilled users get confident mistakes they cannot detect.

That answer contains a claim the rest of this article proves. AI failure is systematic, and the system points directly at experts.

AI fails in two predictable directions

A language model is a prediction engine trained on a snapshot of the past. That single fact produces two distinct failure modes, and nearly every documented AI disaster falls into one of them.

“The expert delta” is our name for the two things a model cannot supply. Knowledge of what changed after its training data was collected, and knowledge of how a specific case differs from the average case in its training data.

The first direction is staleness. The model’s world ends at its training cutoff, so it confidently applies rules the real world has since replaced. The second direction is regression to the average. The model predicts the typical case, so it hands the struggling business the generic advice and hands the unusual situation the standard playbook. The output is fluent in both directions, and the model gives no warning when it crosses into either one.

Look at the famous failures through this lens and they stop looking random. The 2,000+ court decisions involving AI-fabricated citations, tracked by legal researcher Damien Charlotin, involve fake cases in flawless citation format, because flawless format is what the average of past filings looks like. A Stanford RegLab study found even purpose-built legal research tools returned incorrect information on 17% to 34% of queries. Deloitte’s report misstated a court judgment, a fact that lived outside the model’s snapshot. The pattern is the point. AI is weakest exactly where hiring an expert was always strongest, in current knowledge and in judgment about the specific case.

We ran the test in our own field

Instead of taking that on faith, we ran it. We gave a current AI model this one-line prompt in a fresh chat, with no other instructions: “Write the schema markup for a drug and alcohol rehab center’s website… Include everything a rehab website should have for SEO,” naming a fictional Nashville facility, its address, phone number, and three program types. Schema markup is the structured data that tells Google and AI systems what a page is about. It is a core deliverable in our work, and treatment centers are our specialty.

The response was impressive. Ten organized blocks covering the whole site: organization, service pages, condition pages, staff bios, FAQs, breadcrumbs. It used the correct phone number format, three image aspect ratios per Google’s image guidelines, and properly structured opening hours. It looks like the work of a specialist. Then we audited it against Google’s current validation rules and policies, the ones we deploy against every week.

The audit found 14 failures that would block deployment, fail validation, or violate Google policy.

FailureCountFailure direction
Properties Google’s validator rejects on the types used (medicalSpecialty and availableService on MedicalBusiness, provider and indication on MedicalTherapy, an invalid procedureType value, an unsupported credential date field)6Staleness
Schema Google has retired or deprecated (a FAQPage block, retired May 2026, and a sitelinks search box, deprecated October 2024)2Staleness
Date fields in a format that triggers validator warnings, and voice-search selectors pointing at page elements that may not exist2Staleness
Named prescription medications marked up as Drug entities, which triggers Google’s commerce validation and demands fields a treatment page cannot supply1Staleness
A self-assigned star rating and review on the business’s own page, ineligible under the self-serving review policy Google set in 20191Staleness
Missing entity architecture: no unified graph, and cross-references Google’s testing tool cannot resolve across separate blocks2Average case

The second finding is worse. Woven through those 10 blocks were 36 fabricated data points, stated as fact, for a business the model knows nothing about.

What the AI inventedThe details it supplied
A medical staffA Medical Director named Dr. Jane Smith, MD, a Vanderbilt medical degree, board certification in addiction medicine from a named certifying board, and a certification date
A reputationA 4.8-star rating from 127 reviews, plus a five-star testimonial from “Sarah M.” with a quote and a publication date
Clinical program detailsA 3-to-10-day detox length, a PHP schedule of 5 days a week for 6 hours, an IOP schedule of 3 days for 3 hours, and out-of-state admissions
Business claimsFour named insurance companies accepted, financing offered, a toll-free line, a founding year, an employee count, and a founder
Contact and location dataAn email address, a fax number, five social media profile URLs, map coordinates, and two additional cities served

Read that first row again. Asked for markup for a healthcare business, the model invented a credentialed physician, complete with the medical school and the certifying board, because the average rehab website has a Medical Director bio and the model filled the slot. On a page where lives and licensure are at stake, it manufactured the exact trust signals Google and regulators check for.

And the split lands the thesis on its own. The staleness failures exist because Google changed or enforced rules after the training data was collected, so the model still recommends what the old web did. The fabrications exist because the model supplied what the typical site has, ratings, doctors, schedules, insurers, for a facility that has none of them. The polish is the trap. The more professional the wrapper looks, the more an unqualified reviewer trusts the invented facts inside it.

None of this means the tool failed. The tool did its job, which is predicting typical output. Our team catches these in review because we deploy and validate schema across behavioral health sites weekly, and several of these exact failures are written into our internal standards from past deployments. The AI drafted a site’s worth of markup in seconds. The expert delta is the whole difference between that draft and something safe to publish.

The same tool, opposite results

Our test showed what the delta looks like on one task. A 2025 Harvard Business School and UC Berkeley study showed what it does to whole businesses. Researchers gave 640 Kenyan small business owners a GPT-4 advisor over WhatsApp and measured revenue and profit against a control group.

The average effect was zero, and the average was hiding two opposite stories. Entrepreneurs already performing well gained roughly 10% to 15%. Struggling entrepreneurs lost about 8%. Same tool, similar questions, similar answers. The high performers selected advice that fit their specific situation, like the owner who bought a generator for his area’s rolling blackouts. The strugglers took the generic advice, cutting prices and buying ads, which is the average-case answer and made weak businesses weaker.

A 2023 field experiment with 758 BCG consultants found the boundary inside a single job. On 18 tasks within GPT-4’s ability, consultants using it finished 12.2% more work 25.1% faster at higher quality. On one task chosen to sit just beyond its ability, AI users were 19 percentage points less likely to reach the correct answer. The researchers called the boundary the jagged frontier, and the output reads identically on both sides of it. Only someone who knows the domain can tell which side they are standing on.

What happens when companies swap experts for AI?

The corporate record shows the delta compounding at scale. Klarna announced in early 2024 that its AI assistant did the work of 700 customer service agents. By mid-2025 it was hiring humans back, with CEO Sebastian Siemiatkowski admitting the cost focus had produced lower quality, and the company now routes disputes, fraud, and hardship cases to people. MIT’s State of AI in Business 2025 report found 95% of enterprise generative AI pilots delivered no measurable P&L return despite $30 to $40 billion in spending. Gartner has predicted that by 2027, half of the organizations that cut customer service staff over AI will move to rehire.

One more finding explains why companies keep walking into this. A 2025 METR trial found experienced developers using AI took 19% longer on real tasks while estimating the tools had made them 20% faster. METR’s follow-up softened the slowdown, and the perception gap held. People cannot feel the delta. It has to be measured, and measuring it is itself expert work.

What jobs can AI not replace?

Reframed through the delta, the standard list of jobs AI can’t replace stops being a list of job titles and becomes a description of two capabilities.

  • Tracking a changing world. Law after a new ruling, medicine after new evidence, SEO after an algorithm update, any field where this year’s correct answer differs from last year’s. The model’s snapshot ages every day. The expert’s does not.
  • Handling the case in front of you. The patient who presents atypically, the business with an unusual constraint, the legal matter with no clean precedent. The model regresses to the average. The expert reasons from the specifics.

Add the one thing that sits outside prediction entirely, accountability. Deloitte refunded the fee. The sanctioned lawyers paid the fines. A model cannot hold a license, sign an opinion, or answer for an outcome, so in every regulated field a qualified person must stand behind the work no matter what produced the draft.

What are experts actually for now?

Here is the economic version of the argument, and the sentence to remember. AI collapsed the cost of producing work and left the cost of verifying it untouched.

Hiring an expert was always two purchases bundled together, production and verification. You paid someone to do the work and, in the same salary, to know whether the work was right. AI unbundled them. Production is now nearly free, which makes verification the scarce half, and verification is the half that requires the expert delta. Nobody can check output for staleness without current knowledge, or check it against the specific case without situational judgment.

That is why the smart move is the one the data supports. Keep the expert, hand them the tool. The BCG consultants inside the frontier got their 12% and 25% gains precisely because they were qualified to verify what the model produced. Our schema test is the same story in miniature. The draft took seconds, and the 50 catches took expertise our team built across years of deployments. The saved hours are real. They are only worth something because someone qualified spends minutes instead of finding out about the errors from Google, a regulator, or a client.

So the question for a business owner deciding between an expert and a subscription is not which one is cheaper. It is who on your side can tell when the output is wrong. If the answer is nobody, the expert’s salary did not disappear from your budget. It moved downstream, into rework and consequences, and got bigger.

If search visibility is the function you are weighing, that is our lane. Slacker SEO pairs hands-on technical SEO services with AEO and GEO work produced with AI and verified by people who deploy against Google’s live rules every week. Talk to Slacker SEO about your search strategy.

Frequently asked questions

Will AI replace humans?

Not in work that depends on current knowledge, case-specific judgment, or accountability. Research shows AI raises skilled workers’ output on tasks it handles well and degrades results when users cannot catch its errors. The realistic future is experts working faster with AI, with businesses paying for the expertise either way.

What is the expert delta?

The expert delta is the gap between what an AI model can predict and what a current, situationally aware professional knows. It covers two things, what changed after the model’s training data was collected and how a specific case differs from the average case. AI errors concentrate in exactly those two places.

Is it cheaper to use AI instead of hiring an expert?

Usually only on paper. MIT found 95% of enterprise generative AI pilots delivered no measurable financial return, and Klarna reversed its AI-for-agents swap after quality fell. AI cuts the cost of producing work, and someone still has to pay the cost of verifying it, either upfront or after the errors surface.

What jobs will AI not replace?

Jobs built on tracking change, judging specific cases, and answering for outcomes. Doctors, lawyers, skilled trades, therapists, teachers, and experienced operators all do work where the correct answer moves over time and the case in front of them differs from the textbook. AI changes how those roles work far more than whether they exist.

What is overreliance on AI?

Overreliance on AI means accepting its output without independent verification. It is dangerous because models generate wrong answers with the same fluency as right ones. Courts have logged more than 2,000 decisions involving AI-fabricated material since 2023, nearly all from users who trusted output they could not personally check.

How can I tell if AI output is wrong?

You mostly can’t unless you know the subject, which is the core problem. The reliable safeguards are checking claims against primary sources, checking anything time-sensitive against the current rules, and routing AI-assisted work through someone with real domain expertise before it ships.

Method note

The schema test used this prompt, verbatim, in a fresh AI chat session with no custom instructions or other context: “Write the schema markup for a drug and alcohol rehab center’s website. The center is called Riverside Recovery Center, located at 123 Main St, Nashville, TN 37201, phone (615) 555-0142, offering detox, residential treatment, and outpatient programs. Include everything a rehab website should have for SEO.” Riverside Recovery Center is fictional. The full unedited output ran to 10 JSON-LD blocks. We audited it against Google’s published structured data documentation, Google’s review snippet and deprecation announcements, and Rich Results validation behavior as documented in Slacker SEO’s deployment standards. Failure count treats each distinct rule or policy violation once; fabrication count treats each invented data value once. Reproduce it yourself with the same prompt in any current AI chat.

Sources

  1. Deloitte to partially refund Australian government for report with apparent AI-generated errors. Associated Press, October 2025.
  2. The Uneven Impact of Generative AI on Entrepreneurial Performance. Otis, Koning et al., Harvard Business School / UC Berkeley, 2025.
  3. Navigating the Jagged Technological Frontier. Dell’Acqua et al., Harvard Business School working paper 24-013, September 2023.
  4. The GenAI Divide: State of AI in Business 2025. MIT NANDA, July 2025.
  5. AI Hallucination Cases Database. Damien Charlotin, HEC Paris, accessed September 2026.
  6. Hallucinating Law: Legal Mistakes with Large Language Models. Magesh et al., Stanford RegLab, 2024.
  7. Klarna resumes customer service hiring after AI push. Bloomberg / CX Dive, May 2025.
  8. Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity. METR, July 2025, with February 2026 update.
  9. Google Search structured data documentation, FAQ rich results deprecation, May 2026.
  10. Gartner customer service staffing prediction, reported via CX Dive, 2025.

Schema test — audit receipts (internal, archive with the raw output JSON)

Prompt: verbatim one-liner in the article’s method note. Output: 10 JSON-LD blocks (raw file archived separately). Audited September 9, 2026 against Google structured data documentation and Slacker SEO deployment standards.

14 rule / policy / structural failures

  1. medicalSpecialty on MedicalBusiness (Block 1) — validator rejects on this type
  2. availableService on MedicalBusiness (Block 1) — validator rejects on this type
  3. provider on MedicalTherapy (Blocks 1, 4, 5, 6) — unsupported
  4. indication on MedicalTherapy (Blocks 4, 5, 6) — unsupported
  5. procedureType with non-enum values “Detoxification” / “Residential Rehabilitation” / “Outpatient Rehabilitation” — only NoninvasiveProcedure or PercutaneousProcedure validate
  6. validFrom on EducationalOccupationalCredential (Block 10) — unsupported; only credentialCategory, identifier, url, recognizedBy
  7. Date-only datePublished / dateModified / lastReviewed (Block 3) — full ISO 8601 with timezone offset required
  8. Drug entities for Buprenorphine and Naltrexone (Block 4) — triggers commerce validation; standard says omit and describe in MedicalTherapy text
  9. FAQPage block (Block 8) — rich results retired May 7, 2026
  10. SearchAction sitelinks search box (Block 2) — rich result deprecated by Google October 2024
  11. speakable cssSelector “.entry-content”, “.page-title” (Block 3) — unverified selectors; standard is h1
  12. Self-serving aggregateRating + review on the business’s own page (Block 1) — ineligible under Google’s self-serving review policy (2019); values also fabricated
  13. No @graph; @id references left unexpanded across separate blocks — Rich Results Test does not resolve them; standard requires inline expansion
  14. No page-unique @ids on webpage, service, or person nodes — entity architecture absent below the org/website pair

36 fabricated data points

Reputation (5): ratingValue 4.8; reviewCount 127; reviewer “Sarah M.”; review date 2025-11-02; review quote. Medical staff (4): “Dr. Jane Smith, MD” as founder/Medical Director; Vanderbilt University School of Medicine; ABPM board certification; certification date 2015-06-01. Clinical program details (4): detox 3–10 days; PHP 5 days/6 hours; IOP 3 days/3 hours; out-of-state admissions claim. Business claims (8): Aetna; Cigna; BlueCross BlueShield; UnitedHealthcare; financing accepted; toll-free designation; foundingDate 2010; numberOfEmployees 45. Contact & location (12): email; fax; 5 sameAs social URLs; latitude; longitude; areaServed Franklin; areaServed Murfreesboro; homepage datePublished 2024-03-15. Other (3): slogan; alternateName; reviewedBy profile URL.

Post comment.