DigiVisory.com
AI in the Wild

How Companies Actually Use AI: A Primary-Source Field Guide

Jeff EllisAugust 10, 2026 21 min read

Last updated: September 1, 2026

Illustration for How Companies Actually Use AI: A Primary-Source Field Guide

Key takeaways

  • Trust AI claims only when supported by primary sources like engineering blogs and job postings.
  • Engineering blogs provide detailed, peer-reviewed evidence of AI systems in production.
  • Job postings reveal current AI investments and required skills, indicating real projects.
  • Executive and investor statements often exaggerate AI capabilities and should be treated skeptically.
  • Most deployed AI systems are supervised and narrowly scoped, not fully autonomous.

Every week another company announces it is "all in on AI." Every week another vendor tells you its platform will run your business while you sleep. And every week, the gap between what gets announced and what actually runs in production stays wide open.

This guide is not another roundup of press releases. It is a field guide for operators — founders, operators, product leads, and engineers — who need to know what companies are actually doing with AI, not what they are saying they are doing.

The method here is simple and it is the whole point: we only trust what a company discloses through primary sources. Engineering blogs. Job postings. Patent filings. Regulatory disclosures. Product changelogs. Executive and investor statements. Those are the lowest-trust, highest-signal artifacts a company produces, because they are written for engineers, regulators, and investors — not for the marketing page.

Everything in this guide is directional and source-led. The examples are real, disclosed implementations. The patterns are what the evidence shows. The tools are decision trees and ledgers you can actually use.

If you want the marketing version of how companies use AI, close this tab. If you want to know how to read a company's AI claim against what it has actually disclosed, keep reading.


How to Read What a Company Is Actually Doing With AI

Before you can evaluate any single company, you need a method. The method has two parts: where the evidence comes from, and what the evidence actually proves.

The core discipline is this: a company's AI story is only as trustworthy as the artifact that carries it. A press release is a promise. An engineering blog is a receipt. A patent is a claim about what the company believes it owns. A job posting is a statement about what the company is paying people to build right now. A regulatory filing is a statement the company can be sued for.

So the first skill is not reading AI news. It is sorting evidence by how much it costs the company to lie.

Here is the ranking, from most trustworthy to least. We will call them Source Tiers.


Source Tier 1: Engineering Blogs and Technical Write-Ups

What it is: A company's own engineering team publishing how a system was built, what it does, what failed, and what the numbers were.

Why it is the highest-trust tier: An engineering blog is written by the people who built the thing, for an audience of other engineers who will immediately catch any fabrication. It is peer-reviewed by the harshest possible reviewers: the people who have to maintain the system. When Stripe publishes how its Minions agents ship 1,300 pull requests a week, the engineers who have to review those pull requests are reading it. You cannot fake that for long.

What it proves: That a system exists, how it is architected, what it actually does, and — critically — what its failure modes are. Engineering blogs are the only source tier that routinely tells you what didn't work.

What it does not prove: That the system is profitable, that it is the company's strategy, or that it will survive a budget cycle.

Examples in this guide: Stripe's Minions and Knowledge AI Platform, Shopify's River agent, Cloudflare's iMARS stack, Uber's Genie on-call copilot, Atlassian's Rovo and Teamwork Graph.


Source Tier 2: Job Postings and Hiring Patterns

What it is: The roles a company is actively paying to fill.

Why it is high-trust: A job posting is a contract offer. It names the exact skills, tools, and responsibilities the company needs today. You cannot market your way out of a job posting — it is a direct statement of what the company is building and with what stack.

What it proves: Where a company is actually investing. If a company posts for "Agentic AI Engineer" roles requiring LangGraph, CrewAI, and RAG pipelines, that is a statement that it is building agentic systems with those tools. If it posts for "AI Product Manager" roles, it is building AI products. If it posts nothing, it is not building anything.

What it does not prove: That the roles are filled, that the systems work, or that the investment is large relative to the company.

The pattern to watch: The 2025-2026 labor data shows a striking reversal. After years of declining postings in high-AI-exposure occupations, those same roles rebounded sharply starting around February 2025 — and the rebound is concentrated in senior roles. One analysis found 71% of the net increase in software development postings came from senior roles, with 37% of the growth in titles that explicitly mention AI. That is a direct, disclosed signal: companies are hiring senior people to use AI as a force multiplier, not hiring juniors to be replaced by it.


Source Tier 3: Patent Filings

What it is: What a company believes it owns, and what it is willing to spend legal fees to protect.

Why it is medium-high trust: A patent is a legal claim. Filing one costs real money and real time, and the claim is examined. It tells you what a company thinks is defensible and strategically important.

What it proves: Where a company is directing its R&D and what it believes is proprietary. The 2025 patent data is revealing: Samsung led with over 1,000 AI patents, followed by Alphabet and IBM. But the more interesting signal is who else is filing. Capital One ranked among the top AI patent holders globally — a bank, not a tech company. That is a disclosed statement that a financial institution is building proprietary AI, not just buying it.

What it does not prove: That the patented technology is deployed, that it works, or that it matters commercially. Many patents are defensive or aspirational.


Source Tier 4: Regulatory Disclosures and Product Changelogs

What it is: What a company is legally required to say, and what it ships.

Why it is medium trust: Regulatory filings (SEC 10-Ks, AI use case inventories, board disclosures) are statements a company can be penalized for getting wrong. The SEC has been actively pursuing "AI-washing" enforcement — companies making misleading claims about their AI capabilities. That means the risk factors and business descriptions in a 10-K are now a place where companies must be careful not to overclaim.

Product changelogs are different but equally valuable: they are the unglamorous record of what actually shipped, version by version. A changelog entry that says "added RAG grounding to the assistant" is a receipt.

What it proves: That a company has a compliance posture around AI, what it ships over time, and what it is willing to put in writing under legal risk.

What it does not prove: That the disclosed AI is material to the business, or that the changelog features are widely used.


Source Tier 5: Executive and Investor Statements

What it is: What leadership says on earnings calls, in interviews, and in investor communications.

Why it is the lowest-trust tier: Executive statements are the closest thing to marketing that still gets treated as disclosure. They are optimized for narrative, not accuracy. They are the tier where "AI" gets used as a synonym for "growth" and where aspirational products get described as shipped.

What it proves: What leadership wants investors and the market to believe. That is real information — it tells you where the company is positioning itself — but it is not evidence that the AI works.

What it does not prove: Almost anything about the technology itself.

The critical rule: Executive statements are only useful when they are specific and falsifiable. "We are all in on AI" is noise. "Our AI assistant handles two-thirds of customer service chats and performs the workload equivalent of 853 full-time employees" is a claim you can check against other tiers.


What the Pattern Evidence Shows as of 2025-2026

Now we get to the substance. When you sort the disclosed evidence across these tiers, three patterns emerge that cut against the marketing narrative.

Pattern zero, the meta-pattern: The most credible disclosures — the engineering blogs — describe AI as boring, supervised, and narrowly scoped. The least credible disclosures — the executive statements — describe AI as autonomous, transformative, and everywhere. The gap between those two is the single most important thing to understand about how companies actually use AI.

The companies that are actually shipping AI at scale are not the ones making the loudest claims. They are the ones publishing engineering blogs about permission-aware retrieval, deterministic guardrails, and human review gates. That is the anti-consensus finding of this entire guide: the more a company actually uses AI, the more its own engineers describe it as supervised and constrained.

Let's look at the three concrete patterns.


The Autonomy Theater Problem

Before the patterns, you need the lens. Call it the Autonomy Theater Problem.

Autonomy theater is what happens when a company markets AI as fully autonomous while the actual implementation depends on fragile, manual human intervention. It is the gap between the demo and the deployment.

The telltale signs are consistent across the disclosed evidence:

1. The "assistance vs. replacement" test. If a team must perform the task alongside the AI, it is assistance, not autonomy. Real transformation means the task runs without a human in the loop. Most "autonomous" systems fail this test.

2. The "no-login" test. A genuinely autonomous system keeps working if nobody logs in for a week. Theater-based systems stop progressing the moment human attention stops. This is a brutal but fair test, and most enterprise AI fails it.

3. The "debugging confession" test. Listen to what operators actually say. One founder who claimed to run dozens of businesses via AI agents admitted that half their time was spent debugging the setup and another third improving it — leaving only about 20% on productive work. That is not autonomy; that is a very expensive hobby.

4. The "workaround tax" test. Many companies pay for AI that does not eliminate the underlying task — it just moves it to a new interface. The manual work still happens; it is just now called "reviewing the AI's output." If the human effort is unchanged, the AI is theater.

The disclosed evidence is blunt about this. Only a small minority of organizations have achieved what could genuinely be called agentic orchestration in production. The rest are running supervised copilots and calling them autonomous agents. The autonomy theater problem is not a bug in the industry; it is the industry's default marketing mode.

Now the three real patterns.


Pattern 1: Internal Document Processing and Search Is the Workhorse

The single most common, most disclosed, most boring use of AI inside companies is internal document processing and search. This is not the headline-grabbing stuff. It is the workhorse.

What the evidence shows: Companies are not using AI to replace their workforce. They are using it to make their own knowledge findable. The disclosed implementations are almost all the same shape: index the company's documents, make them searchable in natural language, and ground the answers in the source material with citations.

The engineering reality: The engineering blogs are remarkably consistent about what makes this work. It is not the model. It is the retrieval, the permissions, and the grounding.

  • Permission-aware retrieval is non-negotiable. Atlassian's Rovo is built on a "Teamwork Graph" that enforces access control at the point of search — if a user lacks permission to a document, the search engine physically cannot retrieve it. This is the difference between an enterprise tool and a data leak.
  • Hybrid retrieval beats pure vector search. Uber's Genie on-call copilot found that pure semantic search failed on structured data like tables and policy variations. Its fix was an "Enhanced Agentic RAG" combining vector search with keyword (BM25) retrieval, plus LLM agents for query optimization and source identification. The result: a 27% relative increase in acceptable answers and a 60% reduction in incorrect advice.
  • Grounding is the whole game. Atlassian reports that grounding Rovo's responses in the Teamwork Graph improved accuracy by 44% while cutting token usage by 48%. The graph is not a gimmick; it is the difference between a model guessing and a system citing.

The disclosed scale: The customer evidence is concrete. Booking.com deployed Glean as its first company-wide AI platform across 14,000 employees, using it to cut IT ticket resolution time, compress promotional video production from eight weeks to two, and index over 500,000 monthly survey responses. Zillow reported 80% adoption of Glean. Ericsson trained over 20,000 employees on it and deployed over 2,700 agents.

The operator takeaway: If you want to know what "how companies use AI" actually means in practice, it means this: companies are spending real money to make their own documents searchable. It is not glamorous. It is not autonomous. It is supervised, permission-gated, citation-grounded retrieval. And it is the pattern with the most disclosed evidence behind it.


Pattern 2: Customer-Facing AI Is More Supervised Than Advertised

The second pattern is the one that most directly contradicts the marketing. Customer-facing AI — the chatbots and agents that get the press — is more supervised than the headlines suggest.

The Klarna story is the canonical case. In February 2024, Klarna launched an OpenAI-powered customer service assistant. In its first month it handled 2.3 million conversations — two-thirds of the company's customer service volume — and Klarna equated that to the workload of 700 full-time agents. Resolution time dropped from 11 minutes to under 2 minutes. By late 2025, Klarna said the assistant handled the workload equivalent of 853 employees and drove roughly $60 million in annual savings.

That is the headline. Here is the part that gets less attention: Klarna's CEO admitted the company "overpivoted" toward cost-cutting and had to rehire human agents. The company acknowledged that over-reliance on automation hurt service quality on complex, sensitive, and high-empathy cases. By mid-2025, Klarna was hiring an "Uber-type" flexible workforce to handle the cases the AI could not.

The disclosed lesson is not "AI replaces customer service." It is "AI handles the routine tier, and humans handle the edge cases — and the edge cases are where the brand lives."

The supervision pattern is everywhere. Sierra, the enterprise AI agent company serving over 40% of the Fortune 50, sells outcome-based pricing where clients pay for successful resolutions. Its agents are designed to perform real actions — refunds, subscription changes, identity verification — but they operate within strict policy guardrails and hand off to humans. SiriusXM's Harmony agent, built on Sierra, is explicitly positioned as a way to reduce human escalation while keeping humans in the loop for the hard cases.

The disclosed failure cases are just as instructive. The Commonwealth Bank of Australia reversed planned job cuts after its AI voice bot failed to reduce call volumes as expected. That is a disclosed, real-world counterexample to the "AI replaces everything" narrative — and it is exactly the kind of evidence you only get from primary sources.

The operator takeaway: Customer-facing AI is real, it is scaling, and it is handling a meaningful share of routine volume. But the disclosed evidence shows it is supervised, policy-constrained, and paired with human escalation paths. The companies that treat it as a full replacement are the ones that end up rehiring.


Pattern 3: Most Transformation Is AI-Augmented, Not Autonomous

The third pattern is the deepest and the most anti-consensus: most of the transformation happening inside companies is AI-augmented work, not autonomous work.

The engineering evidence is overwhelming. Look at what the most credible disclosures actually describe:

  • Stripe's Minions ship over 1,300 pull requests a week. But the architecture is a "blueprint" that alternates deterministic nodes (linters, test suites, parsing) with agentic nodes (LLM reasoning). The agents have submission authority, not merge authority. Every pull request is reviewed by a human. This is not autonomy; it is a highly engineered human-in-the-loop system.
  • Shopify's River co-authors about one in eight merged pull requests. But it operates only in public Slack channels, and every interaction is visible and reviewable. Shopify's own framing is that the transparency is the product — it builds "AI taste" across the organization. That is augmentation, deliberately designed to be supervised.
  • Cloudflare's iMARS stack reached 93% adoption across R&D, routing 47.95 million AI requests in a 30-day period. But the architecture has an explicit "enforcement layer" — an AI code reviewer that analyzes 100% of merge requests against codified internal standards. The whole point is to keep AI output inside human-defined guardrails.
  • Google's disclosed milestone that over 25% of new internal code is AI-generated is real — but it is framed as AI-augmented engineering, with human oversight as the "crucial layer."

The workflow-redesign evidence is the sharpest. A landmark field experiment across 515 startups found that firms which redesigned their production processes around AI — reorganizing or eliminating chains of work — achieved roughly 90% higher revenue than firms that used AI only to speed up individual tasks. The difference is not the AI. The difference is whether the company treats AI as a way to do the same work faster (augmentation) or as a reason to change what the work is (transformation).

The operator takeaway: The companies getting real value are not the ones chasing full autonomy. They are the ones that put AI inside a supervised workflow, redesigned the workflow around it, and kept a human accountable for the output. Autonomy is the marketing. Augmentation is the reality.


Pattern 4: The Serious Deployments Are Building Memory, Not Just Calling Models

The fourth pattern is the least discussed and the most predictive, and you can read it directly out of the primary sources if you know what to look for.

Companies whose AI actually works in production are not distinguished by which model they call. They are distinguished by what they built underneath it.

What the evidence looks like

Engineering blogs from teams running real production AI spend most of their words on things that are not the model:

  • how documents are chunked, versioned, and kept current,
  • how permissions are enforced at retrieval time rather than at the answer,
  • how feedback and corrections are captured and fed back in,
  • how evaluation sets were built and how they are maintained,
  • how the system handles a source of truth that changed yesterday.

Job postings say the same thing in a different register. A company hiring data engineers, evaluation engineers, and knowledge or ontology specialists alongside its machine learning roles is building a substrate. A company hiring only prompt engineers is building a feature.

Patent filings are the clearest tell of all. Very few of the AI patents filed by operating companies claim a model. They claim retrieval methods, indexing schemes, context assembly, provenance tracking, and human-in-the-loop routing — the surrounding machinery. Companies patent what they think is defensible, and they do not think the model is defensible, because they are renting it from the same three providers as everyone else.

Why this pattern predicts durability

Models are a rented capability. Every company in the market has access to roughly the same frontier, on roughly the same schedule, at roughly the same price. A capability that arrives as a platform feature is table stakes the day it ships — useful, necessary, and not an advantage, because a competitor can buy it tomorrow.

What cannot be bought on that schedule is a company's own operational memory: its exceptions, its policy interpretations, its accumulated corrections, the data its own operations generate. That is the part the disclosures show serious teams investing in, and it is why their deployments survive a model transition instead of restarting at it.

Reading the pattern in a specific company

Signal in the evidence Reading
Engineering blog is mostly about retrieval, evaluation, and data freshness Building a substrate — durable
Blog is mostly about which model and what prompt Building a feature — fragile
Hiring data, evaluation, and knowledge roles Investing in memory
Hiring only prompt and integration roles Consuming a platform
Patents claim retrieval, provenance, routing Believes the substrate is the asset
No patents, no engineering disclosure, heavy press Insufficient evidence — treat as marketing
Documented human review with correction capture Learning is being retained
Documented human review with no correction capture Paying for oversight, keeping none of it

That last pair is worth dwelling on. Plenty of companies disclose a human review step — it is the standard answer to a safety question. Far fewer disclose what happens to the reviewer's edit. A review step that gates output is a cost. A review step that captures the correction is an investment, and over a few years it is the difference between a program that compounds and one that pays the same build cost forever.

The uncomfortable implication for benchmarking

The common way to answer "how do companies use AI" is to survey what the leaders in your industry deployed and match it. The evidence suggests that produces the same capability everyone else has, on the same schedule, with none of the underlying substrate that made it work for the company you copied.

Copying the deployment gets you the visible layer. The part that mattered — the memory, the evaluation discipline, the correction loop — does not travel, because it was specific to that company's operations. Read the disclosures for the method, not the menu.

How to Evaluate a Company's AI Claim Against Primary Evidence

Here is the practical tool. When you hear a company claim it is "AI-powered" or "autonomous" or "transforming the industry," run it through this decision tree. It will tell you, in about two minutes, whether the claim is backed by evidence or is theater.

Step 1: Identify the claim's source tier.

  • Is it an engineering blog? High trust. Proceed.
  • Is it a job posting? High trust. Proceed.
  • Is it a patent? Medium-high trust. Proceed with caution.
  • Is it a regulatory filing or changelog? Medium trust. Proceed.
  • Is it an executive statement? Lowest trust. Demand corroboration from a higher tier before believing anything.

Step 2: Ask the falsifiability question. Can the claim be checked? "We are all in on AI" is not falsifiable. "Our assistant handles two-thirds of chats" is. If the claim cannot be checked, it is not evidence — it is narrative.

Step 3: Apply the autonomy tests.

  • Assistance or replacement? If a human must do the task alongside the AI, it is not autonomous.
  • Does it work if nobody logs in for a week? If not, it is theater.
  • Is the human effort actually reduced, or just relabeled as "reviewing the AI"?

Step 4: Look for the supervision disclosure. The most credible companies voluntarily disclose their guardrails: human review gates, permission-aware retrieval, deterministic checkpoints, escalation paths. If a company describes its AI as fully autonomous with no supervision, that is a red flag, not a selling point. The companies actually shipping AI are the ones telling you how they constrain it.

Step 5: Check for the boring details. Does the company disclose retrieval architecture, permission models, evaluation frameworks, or failure modes? The presence of boring engineering detail is the single best predictor that the AI is real. The absence of it is the single best predictor that it is theater.


Evidence Ledger Template

Use this ledger to track any company's AI claims against its disclosed evidence. Fill in one row per claim. This is the artifact that turns "I read about a company using AI" into "I have verified what a company is actually doing."

Claim Source Tier Source (URL / doc) Falsifiable? Autonomy Test Result Supervision Disclosed? Verdict
"Our AI handles 2/3 of support chats" Tier 5 (exec) Q3 earnings call Yes Assistance (humans rehired) Yes (escalation path) Partially verified
"Our agents ship 1,300 PRs/week" Tier 1 (eng blog) Stripe engineering blog Yes Augmented (human review) Yes (merge authority) Verified
"We are all in on AI" Tier 5 (exec) Press release No N/A No Unverifiable / noise
"AI is our core strategy" Tier 2 (job postings) Careers page Yes N/A N/A Verify against roles

How to use it: For every claim, force yourself to name the source tier. If you cannot name a tier higher than Tier 5, the claim is unverified. If the autonomy test fails, the claim is theater. If no supervision is disclosed, treat the claim with suspicion — the real systems all disclose supervision.


FAQ

Q1: What is the single most common way companies actually use AI?

Internal document processing and search. The disclosed evidence is overwhelming: companies are spending real money to index their own knowledge and make it searchable in natural language, with permission-aware retrieval and citation-grounded answers. It is not glamorous, it is not autonomous, and it is the pattern with the most engineering-blog evidence behind it. If you want to know how companies use AI, start here.

Q2: Is customer service AI actually replacing human agents?

Partially, and the disclosed evidence is more nuanced than the headlines. Klarna's assistant handles a large share of routine chats, but the company's own CEO admitted it "overpivoted" and had to rehire humans for complex, sensitive cases. The pattern across the evidence is a hybrid: AI handles the routine tier, humans handle the edge cases, and the edge cases are where the brand lives. Treat "AI replaces customer service" as marketing, not fact.

Q3: How can I tell if a company's AI claim is real or just marketing?

Run it through the evidence tiers. An engineering blog is high trust; an executive statement is low trust. Ask if the claim is falsifiable — can you check it? Apply the autonomy tests: is it assistance or replacement, does it work without human login, is human effort actually reduced? And look for the supervision disclosure: the most credible companies voluntarily describe their guardrails. If a company describes fully autonomous AI with no supervision, that is a red flag.

Q4: What is the "autonomy theater problem"?

It is the gap between how AI is marketed and how it actually runs. Companies market AI as fully autonomous while the real implementation depends on fragile, manual human intervention. The telltale signs: a team must perform the task alongside the AI (assistance, not replacement), the system stops progressing if nobody logs in, and the human effort is relabeled as "reviewing the AI" rather than eliminated. Most "autonomous" enterprise AI fails these tests.

Q5: Is most AI transformation autonomous or augmented?

Augmented. The most credible engineering disclosures — Stripe's Minions, Shopify's River, Cloudflare's iMARS — all describe AI as a supervised, human-in-the-loop system with explicit guardrails. The field-experiment evidence shows the companies getting real value are the ones that redesign workflows around AI while keeping a human accountable for the output. Autonomy is the marketing; augmentation is the reality.

Q6: What is the best single source tier to trust?

Engineering blogs and technical write-ups. They are written by the people who built the system, for an audience of engineers who will catch any fabrication. They are the only tier that routinely tells you what didn't work. If a company publishes how a system was built, what it does, and what its failure modes are, that is the highest-trust evidence available.

Q7: How should I evaluate a company's AI claim before investing or partnering?

Use the Evidence Ledger. Name the source tier for every claim. Demand falsifiability. Apply the autonomy tests. Look for disclosed supervision. And check for the boring engineering details — retrieval architecture, permission models, evaluation frameworks, failure modes. The presence of boring detail is the best predictor that the AI is real. The absence of it is the best predictor that it is theater.

Jeff Ellis

Writing at DigiVisory.com. Practical AI education for operators.

Get articles like this in your inbox.