The Hypothesis Company
Stephen Messer · Co-founder, Collective[i] & Intelligence.com. Co-founder, LinkShare (sold to Rakuten, $425M). EY Entrepreneur of the Year and 2x Deloitte Fast 50 Winner. Board Member, Spire Global (NYSE: SPIR). Join Intelligence.com
Harvey just closed at $15.5 billion. Its ARR quadrupled from $100 million to more than $400 million in thirteen months. Legora crossed $100 million in the same window from a base of $3 million. Both sell AI to lawyers. Both crossed those numbers because 80 percent of the AmLaw 100, plus more than 800 firms across Europe, moved to them faster than most enterprise buying committees can approve a $100,000 software tool.
Your general counsel already bought a seat. Your outside counsel is running on it. Your board has not been briefed. Meanwhile, one of your own teams is six months into evaluating a tool nobody has tested, at the third meeting called to discuss whether to schedule the fourth. The gap between those two rooms is the single biggest predictor of who wins from here. It is bigger than model choice. It is bigger than budget. It is bigger than industry.
On one side, companies that treat every question about AI as a hypothesis to test. On the other, companies that treat every question about AI as a belief to defend. The gap between them was theoretical two years ago. It is a moat now, and the moat is widening every quarter.
This is not a productivity playbook. This is the management philosophy of the AI-first world. The AI-first company is not the old company with better tools. It is a different operating system, and installing it means rebuilding how decisions get made, how teams get organized, and how failure gets treated. Some of the most valuable companies on earth are already running it. Most of yours are not.
I have watched this exact transition happen once before.
I Watched This Happen Once Already.
At LinkShare I had a front-row seat. The brand marketers I met with between 1998 and 2002 were experienced, well-paid, well-respected people. They had spent years buying upfronts, hundreds of millions of dollars committed once a year, and their judgment about which buys would work had been refined over decades of campaigns. They came to my office to learn about affiliate marketing. They also came to confirm their existing view. Do you think your affiliates are hurting our brand equity. Do you think they are stealing revenue that would come to us anyway. These were closing arguments dressed as diligence. Confirm my world and move on.
Within two years every one of those people was gone. What changed was speed. Affiliate marketing let you test a campaign with real money, real traffic, and real results in a fraction of the time an upfront buy took to execute and measure. The leaders who tried it could see in weeks what worked. The leaders who debated it spent those same weeks building arguments for the world they already believed in. The market did not wait.
Then something deeper broke. The brand marketers who had spent years making sure the logo and the colors were perfect on every business card watched Google change its logo every single day. A company that spent nothing on traditional branding and violated every rule in the brand handbook was at the top of global brand lists inside two years. Facebook did it. Amazon did it. The beliefs these marketers built their careers on were correct in the old world. They were irrelevant in the new one. Marketing today operates more like a drug discovery pipeline, running dozens of experiments and killing the ones that fail fast, than like the brand guardians of twenty years ago.
The people who tested moved forward. The people who defended got replaced. That was 2002. The pattern is the same in 2026. It is just running through a different industry.
Law Firms Are Learning This Right Now.
The legal industry is living through the same transition, and the split between the believers and the testers is now visible in the numbers.
A few lawyers got caught submitting AI-generated citations that turned out to be fabricated, and those stories went everywhere. For the traditional law firm, the stories were confirmation. AI is not ready. We were right to wait. The partners who never tried the tools pointed to the hallucination problem as proof their instinct to hold back was justified. They are also the ones pulling their firms further behind every quarter.
Now stop and think about that for a second. The most risk-averse industry on the planet, one that trained itself for 200 years to say no first and yes after seven partner meetings, went from fear of AI to production reliance in under three years. Every hallucination story hit their inbox first. Every liability concern was theirs to eat. They moved anyway.
The numbers say the rest.
Harvey, founded in 2022 by a former O'Melveny litigator and a former Google DeepMind researcher, closed a $550 million round on September 9, 2026 at a $15.5 billion valuation, co-led by Lightspeed and Diffusion. Annual recurring revenue climbed from $100 million in August 2025 to $190 million in January 2026 to over $400 million by September. That is quadrupling in thirteen months. Roughly 80 percent of AmLaw 100 firms use it, alongside more than 3,000 organizations in 60 countries. Total capital raised is over $1.6 billion.
Legora, the Stockholm-based competitor founded in 2023, closed a $550 million Series D in March 2026 at a $5.55 billion valuation, extended in April with NVentures leading. It crossed $100 million ARR in April 2026 and now serves more than 800 law firms and in-house teams across 50 markets, including White & Case, Cleary Gottlieb, Linklaters, and Bird & Bird.
Then there is Kirkland & Ellis, the highest-grossing law firm in the world and the first firm in history to cross $10 billion in revenue. On May 28, 2026 Kirkland announced a $500 million commitment over three to four years to build its own proprietary AI platform, starting with $100 million in 2026 alone. Two hundred and fifty of the firm's lawyers, including 100 partners, are contributing to the design. More than 180 technology professionals inside and outside the firm are building it. The firm still licenses some third-party tools. It just decided that renting the same software every competitor can rent is not enough.
The bigger move sits inside the announcement. Chair Jon Ballis said Kirkland will use the platform to move the firm off the billable hour and onto value-based pricing. Read that sentence twice. The highest-grossing law firm in the world is deliberately walking away from the economic model that produced $10.6 billion in revenue last year, because it decided a bigger one is on the other side. Freshfields is doing the parallel move with Anthropic, building a bespoke legal model with Anthropic's own team. Thomson Reuters shipped its own proprietary legal model in August 2026 on roughly $40 million of engineering spend. The industry that trained itself for 200 years to move last just outran every enterprise buying process on the planet.
Now think about what your own lawyers told you about the risks of AI even a year ago. Every hallucination story. Every liability concern. Every reason to wait a quarter and then another quarter. Their industry is now writing $500 million checks to keep itself from becoming irrelevant.
Do as I say, not as I do. When it comes to saving their own business, lawyers move aggressively. When advising you, they say move slowly. Watch what they do, not what they say. Do not let your legal team risk-manage your company into irrelevance.
Kirkland is not testing whether AI works. It has run those tests and moved on.
In the time it takes your buying committee to decide on a $100,000 software tool, an entire risk-averse industry made the shift. Your outside counsel is running on it. Your competitors' counsel is running on it. Your general counsel is buying it right now. The industry that historically moved last is now moving faster than most of you.
If law firms can do this, the excuse that your industry is different is dead.
We See This Every Single Day.
At Collective[i] this is the conversation we walk into most weeks. A sales leader or revenue operations team will ask, how do we know it will work for us, or what industries are you good with. The same question shows up from the buyout partner running an operating thesis, the VC associate scoring next quarter's pipeline, and the head of pricing at an airline testing a new fare band. Reasonable-sounding. Almost always belief disguised as diligence.
We take 45 minutes to go live, and every client gets two weeks to change their mind at no cost. They could have real data from their own pipeline in less time than it takes to schedule the second meeting where they plan to debate whether to schedule a third. Instead, teams will spend months discussing whether a tool might work, in a category they have never tried, based on no first-hand experience at all. They are debating a hypothesis they could test in an afternoon.
We no longer waste time on that second conversation. Two groups walk into every meeting we take. The first is AI-first or trying to be. They test the tool on their own pipeline, see what breaks and what works, and move fast. We stay in that room. The second is the group we walk out on. They ask reasonable-sounding questions to confirm the answer they already have. We can watch belief killing them in real time. Their competitors take share every quarter. Their best sellers leave for those competitors and beg the new leadership to bring us in, and the losing leadership disbelieves the seller because the seller now works for a rival logo. They cannot get past their belief being wrong. This is happening inside your company whether you believe it or not.
The specific tell is the sentence "we already have AI." What that CEO usually means is one seat on ChatGPT and a Claude subscription. Neither is a strategy. Both are one kind of brain, already a commodity, out of a dozen that matter. I made the substrate-side version of this argument in The Next Computer Is Alive and the board-facing version in The Safest Move You Can Make With AI Will Cost You Everything. Both land in the same place. The gut is running old software. It was optimized for a world where the penalty for a bad bet was a year of wasted work. It has not updated to a world where the penalty for not testing is falling further behind every week.
First Principles Are the Operating System.
There is a specific tell for a hypothesis company. The people running it are willing to say "I do not know yet." I made the deeper case for this in Bubble Talk Is How You Spot Someone Who Missed AI. The willingness to say "I do not know yet" is the tell of someone building conviction. A strong opinion in the first three seconds of a conversation is the tell of someone performing.
First-principles thinking is not a productivity habit. It is the operating system of the hypothesis company. Ask the question the consensus is unwilling to ask. Reason from what has to be true rather than from what the room already believes. Design a test that would tell you if you are wrong. Run it. Update. That is the whole discipline.
The old operating system worked in reverse. Reason from what the room already believes. Design a project that protects the belief. Run it slowly enough that no test ever contradicts it. Update nothing. That is what most companies still run. Most CEOs still call it strategy.
The AI-first management philosophy runs on first principles because it is the only way to break through a belief coalition inside your own company. Every belief your team inherited from the old software era was built on assumptions that no longer hold. If you cannot ask why the assumption exists in the first place, you cannot see what to test.
The willingness to say "I do not know yet" is the tell of someone building conviction. A strong opinion in the first three seconds of a conversation is the tell of someone performing.
Belief Became Politics. Politics Became Culture. Culture Runs the Room.
Belief is never just belief. It is a coalition. The moment a leader holds a belief, the people who share it move toward that leader. The ones who do not either leave or learn to keep quiet. Over time, the people who share the belief work on projects the belief made possible. They win together. They defend each other. They become a team. Then a function. Then a culture. Every mature company is a stack of these cultures. Some hold real signal about what makes the company work. Many are groups of people who share a belief that has stopped being tested and now runs on inertia because the coalition inside the company will not let the belief lose.
That is the political layer. Moving a decision from belief to hypothesis changes more than what happens on Monday morning. It threatens the coalition that formed around the belief being defended. If the hypothesis wins and the belief loses, someone loses status, someone loses budget, someone loses the identity they built inside the company. Those people are rational. They fight back. They fight hard.
Any leader trying to run this playbook Monday morning will hit the coalitions before they hit anything else. The coalition's oldest tactic is exactly the one I diagnosed in The Oldest Trick in Management Just Stopped Working. Propose a project. Buy time. Call it strategy. Delay the test long enough that the belief keeps its status by attrition. The countermove is small and blunt. Name the beliefs. Ask which ones are actually beliefs and which have quietly become the political property of a group that will fight to keep them. Ask where the test threatens someone's identity rather than their opinion. That map is worth more than any strategy deck.
The belief you cannot test is almost always the one someone in the room needs to protect.
This is why AI-first companies flatten. Middle layers are where belief-coalitions live. Flatten the org and you dissolve the coalition before it can vote against the test.
Jensen Huang at NVIDIA runs one of the most extreme versions of this on the planet. Around 60 direct reports. No one-on-ones. Every problem attacked in a group. Jensen said at Stanford Graduate School of Business that CEOs should have the largest number of direct reports because the people reporting to them need the least hand-holding. The number of layers you strip out is roughly the number of coalition points you eliminate. NVIDIA has stripped out about seven. He is not soft on his team. He gives feedback in front of everyone. He tortures people into greatness, in his own words. NVIDIA compounded 70 percent per year for a decade running that operating system.
Bezos at Amazon built a version of the same discipline on two-pizza teams and reversible decisions. The reason has nothing to do with training people to be more open-minded. The tooling and the structure made it faster to test than to argue. Speed beat politics.
The flat org is half the answer. What runs on top of it is the other half.
How the Hypothesis Company Actually Operates.
A hypothesis company asks two questions before anything else. Could this work. And if it does, will the impact on our growth be large enough to matter. If the answer to both is yes, waiting longer to try is the wrong move by definition. You are delaying a potential step change to protect a belief you have not tested.
Amazon calls the operating principle Bias for Action. In the 2015 shareholder letter Bezos argued that most decisions are two-way doors, reversible and safe to walk through quickly. Only a small number are one-way doors that deserve deep deliberation. The problem in most companies is that every decision gets treated like a one-way door. Committees form, review cycles multiply, and the organization slows to a crawl to de-risk something that could have been tested and reversed in a week. In the 2016 letter Bezos put it more plainly. Most decisions should be made with about 70 percent of the information you wish you had. Waiting for 90 percent means you are almost always too slow. Double the number of experiments per year, he added, and you double your inventiveness.
Booking.com runs about 1,000 experiments in parallel at all times, and more than 25,000 in a year. Their design director's rule is one sentence. No HIPPOs. Highest Paid Person's Opinion is not allowed to override a test. Every decision is a democracy. Every decision is tested. Lukas Vermeer, their director of experimentation, has said in public that there are more variants of the Booking website live at any given moment than there have been humans in history. That is a company that treats belief the way it treats a bug.
Netflix killed its own successful recommendation algorithm and replaced it with a test-and-learn framework that tries dozens of approaches and measures every one. Anthropic disclosed in May 2026 that more than 80 percent of the code merged into its production codebase was written by Claude, up from single digits when Claude Code launched. The AI is now writing itself. Mike Krieger, Anthropic's CPO and a Collective[i] advisor, put it in one sentence. Claude is now writing Claude. That is the recursive version of the same discipline. Tests running tests.
These are not eccentric operators. They are among the most valuable companies on earth. They reached that position by replacing belief with experiment at the operating system level.
A hypothesis company treats failure differently. In a belief company, failure is a judgment on the person who believed wrong. In a hypothesis company, a failed experiment is data. Drug companies kill more candidates than they advance. Nobody fires the scientist whose compound did not survive Phase 2. The experiment did its job.
A hypothesis company staffs differently. It promotes people who run good experiments over people who have good instincts. The leader who says "I tested three approaches this quarter and here is what the data showed" is more valuable than the leader who says "I have been in this industry for twenty years and I know what works." One is learning. The other is coasting.
What the AI-First Company Actually Does.
If the hypothesis is the unit of decision, the AI-first company is what you get when the whole operating model is rebuilt around it. There are four moves.
First, the flat org. Every layer of hierarchy is a place where a test can be blocked. Fewer layers, faster tests. Meta cut a layer of managers this year. Google cut 35 percent of managers with fewer than three direct reports. Block cut 40 percent of its staff before Dorsey and Botha published the essay explaining why. Meanwhile, the exact play I diagnosed in The Oldest Trick in Management Just Stopped Working still runs inside most companies. Propose a project. Buy time. Call it strategy. Every manager who runs that play is a manager defending a belief against a test. Every one is a coalition point. Flatten and the coalition dissolves.
Second, the decision to build differently. Most companies apply AI to what they already do. The ones pulling away ask whether what they are doing should exist at all. That is the whole argument of The Art of Subtraction. The sales forecast built on humans typing into a CRM is the cleanest example. AI does not need it. Once the machine is running, the humans no longer build the forecast. The forecast still exists. The job that built it does not. That is the difference between adding AI and being AI-first. Workflows vs. Outcomes made the same point on the tooling side. CRM was built to standardize work. AI needs to know what to optimize for. Bolt AI onto a workflow tool and you get faster workflows. You do not get better outcomes.
Third, the composition of the team. Not headcount. Composition. The AI-first company builds an operating model around a small number of high-judgment humans supported by a large number of digital workers. I walked through the shape of it in Infinite Leverage. Building a Company With More Agents Than People. The constraint is no longer how many people you can hire. It is how clearly you can define what you want intelligence to pursue.
Fourth, the frame. The Companies Winning at AI Are Playing a Different Game walks through the shift in detail. The old game was to accumulate advantages, protect them, and out-execute inside them. The new game is to build faster loops than your competitors and let them run. Loops beat plans. I laid out the operational version in What It Means to Be AI-First. And How to Get There. Not philosophy. What the companies pulling away are actually doing.
Flat structure. Hypothesis substrate. Human plus agent composition. Loop-driven strategy. That is the operating system of the AI-first company. That is the management philosophy for the AI-first world.
Every quarter you spend still running the old operating system is a quarter your best builder gets closer to leaving for one that runs the new one.
The Talent Vote.
Recruiting is where this argument gets personal fast. The best builder on your team, the person who can ship anything you can describe, is choosing between two kinds of company. In one, they run experiments, kill the ones that fail, learn something on every attempt, and try again with what they learned. In the other, they walk into a room where the leader has a gut, and when the gut is wrong, someone in the room takes the blame. Ask any real builder which one they want to work in.
The hypothesis company gives builders room to run. Failure there is a data point rather than a career event. That single fact is why the best people gravitate to hypothesis companies once they have worked inside one. It is also why the belief company slowly loses its best builders to competitors who let them try things. The A-players leave. The B-players stay because the environment matches their risk tolerance. The company gets an average team, argues about the team's average performance, and blames the tools.
Pick the wrong kind of leader and you are not picking an operating model alone. You are picking the talent pool your company will draw from for the next decade. Boards that skip that test in their CEO searches are giving away one of the last real levers they have left. I made the recruiting-side version of this argument in Find Your Builders. Or They'll Leave and Start Without You.
The Choice Is Live Right Now.
The people who tell you they need more time to think about AI are describing a belief. The people who tell you they need to run a few tests to see what actually works are describing an operating system. The gap between those two sentences is the gap between the companies that will still be leaders in five years and the companies that will still be talking about becoming leaders in five years.
At Collective[i] we built the AI that predicts economic outcomes at scale for sales, private equity, venture, and pricing. At Intelligence.com we built the network that connects the people running those companies to each other so the tests move faster than any single organization could run them alone. Both exist because we bet the hypothesis company would win. Every quarter we get more data confirming the bet.
If you are the CEO who says "we already have AI" because someone on your team has a ChatGPT seat, you are the brand marketer in 2004 buying the upfront while Google was changing its logo every day. If you are the partner blocking a Harvey pilot because a lawyer somewhere hallucinated a citation, you are the horse breeder in 1908 explaining why the automobile is a fad. If you are the industry leader still asking whether AI is "worth it," Kirkland's chair just spent one percent of a $10 billion top line answering the same question you have not started asking. Neither of the first two lost because they were stupid. They lost because their belief was the political property of a group inside the company that could not let it lose. Do not be the third.
You cannot fix that by adding a slide to the strategy deck. You fix it by running the test. Give a team two weeks and a real question. Let them try, kill what fails, and keep what works. Do it again next quarter. Do it again the quarter after that. Somewhere between the third and the sixth cycle, the culture will change, because the people who ran the tests will have the answers, and the people who protected the beliefs will not.
That is the whole race. It is running right now. Every quarter you spend on the second meeting to schedule a test is a quarter Harvey adds another AmLaw firm and your competitor's RevOps team stands up another agent that closes deals while your team is still building the deck.
Test something this week.
Build with the operating system, not on top of the old one.
Intelligence.com connects the people running the tests. Collective[i] runs the predictions on your pipeline. Both take less than an hour to try.
Subscribe at reloadnyc.com. Free. No paywall. No course at the end. Just the work.
If this piece changed how you are thinking about how to run your company, forward it to one person who needs to read it. A board member wrestling with a CEO who cannot let a belief lose. A CEO wrestling with a board that thinks a slide is a strategy. A founder trying to explain why the fourth pilot is a bad sign, not a good one. One specific decision-maker in your world. That is how the argument moves.
Post it on X, tag @smesser, and tell me where I am wrong. Share it on LinkedIn if the leadership teams or boards you sit on need to see it, and tag me at linkedin.com/in/stephenmesser. Tag the people already running the operating system. Jensen Huang. Jack Dorsey. Andy Jassy. Brian Armstrong. Sebastian Siemiatkowski. Matthew Prince. The rest of the CEOs on your board list need to see who is not on the list.
Sources
1. Harvey $550M round at $15.5B valuation, closed September 9, 2026, co-led by Lightspeed Venture Partners and Diffusion. ARR $100M (August 2025), $190M (January 2026), over $400M (September 2026). 80% of AmLaw 100 firms; more than 3,000 organizations in 60 countries. Total capital raised over $1.6B. Sources: Bloomberg, SiliconANGLE, PYMNTS.
2. Legora $550M Series D at $5.55B valuation, closed March 10, 2026, led by Accel. $50M April 2026 extension led by NVentures (NVIDIA venture arm). Total funding approximately $866M. 800+ law firms and in-house teams across 50 markets including White & Case, Cleary Gottlieb, Linklaters, Bird & Bird. Sources: Legora press release, Crunchbase News, Crunchbase extension coverage.
3. Jack Dorsey and Roelof Botha, From Hierarchy to Intelligence, essay published March 31, 2026, following Block's ~40% workforce reduction in February 2026.
4. Google management cuts: Brian Welle at Google all-hands, ~35% reduction in managers with three or fewer direct reports over the prior year. Reported CNBC, August 2025.
5. Jeff Bezos, 2015 Amazon shareholder letter (two-way doors / one-way doors framework). Jeff Bezos, 2016 Amazon shareholder letter ("Most decisions should probably be made with somewhere around 70 percent of the information you wish you had"). "Double the number of experiments per year, double your inventiveness" attributed via The Everything Store, Brad Stone.
6. Booking.com experimentation program: ~1,000 concurrent experiments at all times, 25,000+ tests per year. "No HIPPOs" rule attributed to Stuart Frisby, former Director of Design. Lukas Vermeer, Director of Experimentation. Sources: Booking Data Science and ML blog, Marpipe and Good Men Project coverage of Vermeer public interviews.
7. Jensen Huang, ~60 direct reports (fluctuating 36-60 in internal lists), no one-on-one meetings, feedback delivered in group settings. Sources: Lex Fridman Podcast March 2026, Stanford Graduate School of Business talk 2024, Entrepreneur, Fortune. NVIDIA 10-year CAGR ~70%.
8. Netflix test-and-learn framework: Netflix Tech Blog, Product Strategy forum peer review process, cited in Microsoft Research SEAA21 paper on A/B testing at scale.
9. Anthropic recursive development: "When AI Builds Itself," Anthropic Institute paper, June 2026. 80%+ of production codebase written by Claude by May 2026. Mike Krieger public statements.
10. Kirkland & Ellis $500M proprietary AI platform commitment announced May 28, 2026. First law firm to cross $10 billion in revenue ($10.6B). $100M in 2026, $400M more over three to four years. 250 lawyers (100 partners) and 180+ technology professionals inside and outside the firm. Chair Jon Ballis on shift to value-based pricing away from the billable hour. Sources: Reuters via Yahoo Finance, Legal Cheek, PitchBook, Thomson Reuters Institute analysis. Freshfields partnership with Anthropic reported by Reuters, April 2026. Thomson Reuters proprietary legal model announced August 2026, SiliconANGLE.