“Silicon Valley is LLM-pilled”: Saining Xie on why world models need the world
The co-creator of Diffusion Transformers left the LLM race to co-found AMI Labs with Yann LeCun—$1.03 billion raised before any product, and no office in Silicon Valley.
In 60 seconds
- AMI Labs: about 25 people, a $1.03 billion seed round at a $3.5 billion pre-money valuation—described as Europe’s largest seed—and offices in Paris (HQ), New York, Montreal and Singapore. Not Silicon Valley.
- Xie calls AMI a “reverse OpenAI”: its data cannot be downloaded from the internet, so it wants partners that own real-world data—farms, hospitals, factories.
- He argues large language models are “anti–Bitter Lesson”: language is a human-made shortcut, and he worries it is already “polluting” vision.
- His two near-term outlets for world models: always-on AI glasses (a real personal assistant needs one) and robotics—solved “without building robots.”
- He turned down OpenAI in 2018 (Ilya Sutskever called, annoyed) and Ilya’s SSI in 2024.
Saining Xie (谢赛宁), born 1990, studied at Shanghai Jiao Tong University and UC San Diego, teaches at NYU, and spent four years at Meta FAIR and a stint at Google DeepMind. He co-created Diffusion Transformers (DiT), the architecture behind many of today’s image and video generators. He is co-founder and chief science officer of AMI Labs, with Yann LeCun. This was his first long interview.
When one of the people who shaped modern image and video generation says the industry is chasing the wrong target, it is worth knowing his argument—whether or not you agree.
Five ideas worth your time
His case: the benchmark race decides where money goes, so frontier labs no longer define new problems. Researchers who want to work on video understanding end up assigned to caption data for video generators. AMI’s answer is an open, research-first company with a neutral, international face (LeCun is both American and French). He compares it to Mastercard: when Bank of America’s Visa dominated, local banks formed an alliance. World models, he says, are naturally more decentralized.
The idea is old—psychologist Kenneth Craik described it in 1943, control engineers used it for moon landings (model predictive control), and Richard Sutton’s Dyna paper built on it. What matters is the representation: a useful “state” keeps what decisions need and throws the rest away, the way fluid dynamics replaced modelling every molecule.
Sutton’s Bitter Lesson says general methods plus compute beat clever human design. Xie argues language is the cleverest human design of all, so LLM scaling laws mix “knowledge compression” with real understanding. He calls language a crutch—even “opium”—and says the contamination of vision by language is already happening. That puts him openly at odds with Kimi founder Yang Zhilin, who worries vision makes language models dumber.
A four-year-old has seen more visual data than all the text used to train LLMs, he says. Near-term uses: AI glasses that are always on (a real personal assistant, in his view, needs a world model) and robots—he wants to solve the robot “brain” without building hardware, calling it the second half of pre-training.
Following Sutton, he argues that a squirrel surviving in the real world is harder than winning math olympiads. He says most of Silicon Valley doesn’t believe in AMI and most of the rest of the world does. He quotes Jürgen Klopp—“I’m the normal one”—and says he wants to be the team’s battery.
“Silicon Valley is very LLM-pilled.”— Saining Xie
“The world needs World Models. World Models need the world.”— Saining Xie
What it means for you
- If you sit on real-world data—operations, sensors, video of work—world-model labs want partners, not just customers. That is leverage.
- Today’s personal agents (OpenAI dots, Manus Cue) run on language models. Xie’s bet is that a truly always-on assistant needs a world model. If he is right, the agent race has a second round.
- Don’t build where the benchmark race is decided by capital. Pick a problem the frontier labs have stopped defining.
- “Neo labs” are raising billion-dollar seeds before any product (AMI $1.03B; Thinking Machines $2B). The underwriting is a research breakthrough on a multi-year horizon.
- Xie describes a belief split: skeptics in Silicon Valley, believers elsewhere. Geography of capital may matter for these bets.
- Watch the glasses and robotics outlets—those are where a world model would first show revenue.
- Video understanding is, by his account, under-resourced across both academia and industry—an opening for focused teams.
- Expect scaling laws for world models to look different from LLMs: less memorization, more filtering and organizing of information.
- His warning on the “middle-paper trap”: good results without the compute to turn them into breakthroughs.
Read it chapter by chapter
Every argument, example and number from the original, retold in our words in the original order. Short quotes are attributed.
- 01 · “Silicon Valley has been hypnotized”
- 02 · An invisible world beyond language models
- 03 · Language models predict the next token; world models predict the next state
- 04 · Why he calls LLMs “anti–Bitter Lesson”
- 05 · Turning down Ilya Sutskever—twice
- 06 · “Language is opium”—his worry about vision
- 07 · An underdog under industry pressure
- 08 · “Arrogant humans”
- 09 · “42”
Asked whether he would have started a company if Yann LeCun had stayed at Meta, Xie says probably yes—after some agonizing—and that he still doesn’t know whether he would want to be CEO. What drew him in is that AMI’s agenda is exactly what he had wanted to work on. He describes it as a “reverse OpenAI.”
The forward version, in his telling: take the internet as your data source, download it, train a GPT with Transformers, and push the resulting intelligence to market as consumer or business apps. (He adds that he considers “AGI” a completely false concept.) The reverse version starts from the opposite constraint. The data a world model needs cannot be downloaded; there is no shortcut, and no one can walk that road alone. His line: “The world needs world models. World models need the world.”
So AMI imagines something like a grassroots alliance: companies that feel the AI FOMO, have concrete problems and sit on large amounts of real-world data, co-building an initial world model with AMI. They use it to create value, generate more data, and that data improves the model—a closed loop.
Why four offices from day one—Paris as headquarters, plus New York, Montreal and Singapore? LeCun’s standing gives the company a neutral, international face; he is both American and French, and not being in Silicon Valley makes it easier to attract partners around the world. Xie tells a story a mentor told him: Bank of America launched the first consumer card that became Visa, profited quietly, and dominated; smaller local banks, unable to compete alone, formed an alliance and launched Mastercard. He is careful to say AMI isn’t copying that model, but that world models are a more decentralized story that naturally resists monopoly.
That is also where AMI’s openness comes from. It is a serious startup and won’t open-source everything, but it wants to sit between a pure research lab and a fully closed model company. He sees the same balance in himself: not a senior, established professor, and not an 18-year-old who can move into a Shenzhen factory to collect data. And yes, skipping Silicon Valley was deliberate. The Valley is “very LLM-pilled,” he says—but hypnotized people eventually wake up, and AMI may open an office there when they do.
The decision to start a company was, he admits, a gut call. Mentors from the Bay Area—investors and founders—told him academia had a hard ceiling. The problem is compute. Academia gave him room to find his direction, but he felt he was nearing a “middle-paper trap,” like the middle-income trap: publishing decent papers without the resources to turn ideas into breakthroughs.
Around the autumn of 2025 a mentor suggested he ask LeCun, who seemed unhappy at Meta—this was before Meta’s turbulence, the arrival of Alexandr Wang and the FAIR layoffs. Xie’s first reaction was that it was impossible: LeCun was a godfather of AI and a pure researcher. The next Monday, in a scheduled one-on-one, LeCun told him first: keep it quiet, but he had decided to leave and start a company. When Xie asked about the business model, it matched what he had imagined almost exactly—something no company in any country, Bay Area included, could do today: a research-heavy “world model” company that isn’t under pressure to launch a product and earn revenue immediately, yet isn’t an academic lab either.
Why not join an existing frontier lab? “Closed,” he explains, means no open source, no papers, not even named blog posts. At Google DeepMind he was the only person in the whole generative AI organization with a joint academic appointment. Big labs look down on academic work, publish nothing, and even inside one company the research and product sides barely talk. The core model teams have one goal: stay at the front of a highly competitive race. That squeezes out the oxygen research needs.
He lays out the value chain as he sees it: at the top sit narratives—the Bitter Lesson, AGI, LLMs. Narratives define benchmarks; benchmarks decide where resources go; and resource allocation drifts away from what researchers think is right. His example: everyone agrees video understanding matters and is unsolved, yet capable researchers at every company end up assigned to write captions for video-generation training data, because that is the only role connected to the value chain. In this finite game, companies have lost the ability to define problems—something early OpenAI had, with GPT and CLIP. His answer is simply to escape, and to build a more researcher-friendly organization.
AMI’s early team came from OpenAI, DeepMind and xAI, he says, motivated by research rather than money or an IPO. He worries the industry has swung too far toward “lower everybody’s ego”: people become team members, but also replaceable screws in a giant machine. He wants young researchers to have visibility and a career arc, and says he doesn’t believe in assembling famous “superhero” researchers and hoping for chemistry.
What LeCun told him, as Xie understood it: world models are the intelligence the real world needs. Outside Silicon Valley and outside the LLM narrative there is an invisible world—farms, hospitals—full of physical problems LLMs can’t solve, and people anxious that they won’t even get a seat at the table. That world is invisible in the Valley’s story but is a huge market. And world models need it: problems should come from real needs and industrial production, and the data—including vast amounts of continuous, high-dimensional, noisy non-visual signals—never gets uploaded to YouTube, whose data skews toward entertainment.
Xie is frank that he doesn’t understand business and has never run a startup, which makes him both anxious and fearless. The company’s most important product, he says, is a research breakthrough—something at least on the scale of the Transformer or ChatGPT. That’s why AMI also has a CEO, serial entrepreneur Alex LeBrun, while Xie keeps the chief science officer title he loves. He describes LeCun as a teenager who never stopped at 65: model airplanes, astrophotography, electronic music and jazz (his website lists New York jazz clubs), and sailing. When Xie named a March paper “Solaris” after the novel and Tarkovsky film, LeCun asked which adaptation he meant—1972 or 2002. Xie says he didn’t agonize long over joining: LeCun talks like someone casting spells.
His definition is deliberately plain. Take a system or environment in some state; apply an action or intervention; learn a function that, given the current state and the action, predicts the next state. It isn’t even new. In 1943 the psychologist Kenneth Craik proposed that people carry such a model in their heads, letting them foresee the consequences of actions and decide what to do. Control engineers built on the same idea for moon missions; model predictive control repeatedly samples action sequences, rolls them forward with a model, scores them with a cost function, executes the first step of the best one and repeats.
Richard Sutton’s Dyna paper made the same point for reinforcement learning, contrasting reactive policies with model-based ones—roughly System 1 versus System 2. Sutton, Xie notes, called pure reinforcement learning primitive precisely because it lacks a world model. With a good enough world model you get planning, which he treats as close kin to what the LLM world calls reasoning.
So it is “given your action, predict the next state.” The state should be a sufficient statistic: keep what prediction and decisions need and ignore the rest. You could describe a room down to every texture and sound wave, but an agent that only wants to chat needs a few facts. Learning such states is representation learning—hierarchical, increasingly abstract, increasingly useful for decisions. We model airplanes with fluid dynamics and the Navier–Stokes equations, not by simulating every molecule; more abstraction lets you describe a larger world.
Language, he says, is a packaged, highly condensed abstraction that human society already built. AMI wants a different abstraction: a latent representation that people can probe but that isn’t bound by the syntax of language. Hence his provocative claim that LLMs are not a triumph of Sutton’s Bitter Lesson but its opposite: the lesson says to strip out human cleverness and let search and learning find answers, and language is the cleverest human structure of all. Language will matter in future systems, but chain-of-thought and the rest of the LLM toolkit are, to him, transitional. He cites research suggesting reasoning traces can be post-hoc, that wrong chains still reach right answers, and that gains may come from simply generating more text.
He argues LLMs have structural safety and controllability problems because they lack a true world model; alignment today means feeding models huge amounts of data about what not to say. A real world model could predict the consequences of an action and avoid harm at inference time—how do you make sure a knife-wielding kitchen robot doesn’t turn and hurt you?
On scaling, he agrees that compression is intelligence and that bigger models generalize better. But he thinks LLM scaling laws are “watered down”: benchmarks reward retrieving facts, so knowledge compression and world understanding get mixed together. World models should scale differently—less memorization, more filtering and organizing, like people. Human senses take in something like a billion bits per second while speech and action run at 10–100 bits per second; the brain’s job is to filter.
LLMs are a crucial part of an intelligent system, not the whole. Turning your head a few degrees produces hundreds of frames; tokenizing each frame into a long flat sequence makes no sense, and Transformers pay equal attention to every token. A world model needs physical understanding, large associative memory, reasoning and planning, counterfactual and causal inference, and controllability. It isn’t a replacement: without LLMs we couldn’t even talk about world models. But language is a communication tool used with intent; LLMs look more like an extension of search engines.
Recent video generation moved things a step: language shifts from being the modeled object to scaffolding for prompts, and models start assigning probabilities to raw pixels—learning, say, why four-legged cats are more common than three-legged ones. Even pixels aren’t fully Bitter Lesson, he adds; they are a human-made grid. Showing output to people is an interface, not the core.
How do you train it without an internet of text? His biggest bet: the era of “downloading the internet” gives way to “downloading humans.” A four-year-old has seen more video than all the tokens used to train LLMs. Video is the first step—he notes that roughly 30 minutes of YouTube uploads would be plenty, terms of service permitting. Near-term outlets: AI glasses, because a real personal assistant must be always on and needs a world model (he compares it to Whoop or Oura, which make decisions from far too few signals), and robotics, whose real problem is the “brain.” He isn’t fond of the term “world model”—it risks becoming a hype bucket—but likes one professor’s joke that it’s a model of the world, not of words.
In 2018 he interviewed at OpenAI: five or six hours alone in a small room on one problem, handwritten in pencil on a sheet of A4 by John Schulman. He got the offer and declined without discussion, because Kaiming He, Piotr Dollár and Ross Girshick—the leading computer-vision trio—were at FAIR. Ilya Sutskever called, sternly asking whether it was about money. Xie can’t recall the figure; top PhDs then commanded roughly $400,000–500,000, at least triple that now, and OpenAI’s offer wasn’t low.
The second call came in July 2024, when Ilya had just founded SSI. This time Xie had just started at NYU. They talked less about pay than about how to give future AI the capacity for love. Xie asked how Ilya saw multimodality and perception; Ilya thought it was largely solved. Xie concluded SSI’s language-first path wasn’t what he wanted to design—not a rift, he insists, just different people doing different things. On AI and love, their conclusion was only that it matters: without it the future is uncertain and dangerous. He asks why we trust our children but fear a new kind of intelligence—and says part of the answer should be technical: making AI more trustworthy and controllable, one reason to build world models.
He isn’t discouraged that computer vision was pushed to the margins after ChatGPT; he thanks LLMs, because they let vision grow into real multimodal intelligence. He draws the field’s history along an axis from simple tasks—MNIST digits, 32×32 CIFAR images, ImageNet—to structured detection and segmentation, and then to multimodal learning, where language becomes a flexible interface. The benefit is freedom to pose any question; the risk is that many “multimodal” tasks become pure language problems. He sees that as a huge opportunity: once AI must deal with real tasks in the real world, weak visual representations become a major flaw. LeCun’s image: we are walking on a language-model crutch—able to walk, not to run in the Olympics.
Real intelligence, for Xie, means interacting with the real world. LLMs mostly work in digital space—memorizing facts, legal advice, education—while vision must handle continuous, high-dimensional, noisy domains: industrial process control, sensor modeling, anything where you need to predict how a system responds to an intervention. He puts LLMs acting through code on one end and general-purpose robots with their own brains on the other; the future of visual and multimodal intelligence lies in between. He wants to solve robotics “without building robots”: hardware is advancing fast (he cites Unitree’s robots on the Spring Festival gala), but someone must focus on pre-training the robot brain.
Why did language get a scaling law before vision? Vision still has no real one, he says; video diffusion shows some scaling behavior, but the balance of parameters and data differs. His contrarian view: language modeling isn’t really self-supervised but strongly supervised—millennia of human processing stored as tokens and uploaded to the internet; free to researchers but full of labels. Language describes only the output side of the world and is a communication tool, not a thinking or decision tool; it trades away dynamics (“the cup broke,” not how it broke).
This puts him at odds with Kimi founder Yang Zhilin, who worried about a “dumb multimodal” model and adding vision that might lower intelligence. Xie agrees you don’t want a dumb multimodal model—but says without vision a model will certainly be dumb, and invokes Moravec’s paradox: what is easy for people is hard for machines. His own worry runs the other way. Language is a poison, or opium: more of it always feels better, but it is a shortcut; lean on a crutch forever and your leg muscles never develop. The pollution of vision by language, he says, is already happening.
World models lack a shared definition because they are a goal, not a technical route; everyone, LLM or video diffusion, is heading there. Video generators—Sora, ByteDance’s models, Genie, Runway, Luma—focus on world simulators that render consistent, beautiful video. Fei-Fei Li’s World Labs emphasizes explicit 3D representations as an interface for collaboration and creation that can guarantee spatial correctness, which a context-limited generative simulator can’t. AMI wants a predictive brain that raises intelligence itself. He calls the routes complementary—you could train agents inside World Labs’ explicit spaces.
To the investor wisdom that “silver spoon” startups never succeed, he pushes back: AMI is a grassroots alliance. LeCun is no grassroots figure, but among investors and the industry he is half-supported, half-opposed—a man still pursuing something not yet proven. They are an underdog surviving under industry pressure; their funding, however large, is tiny next to what LLMs command. Raising money was easy with LeCun, he admits, but they still have to deliver a breakthrough. Most of Silicon Valley doesn’t believe in them; most of the rest of the world does. He’s all in.
He agrees with LeCun that AGI is a false premise. In a debate with Demis Hassabis, LeCun argued that with two million optic nerve fibers the space of possible visual functions is astronomically large, yet humans can process almost none of it; human intelligence is highly specialized. Reading Frans de Waal’s Are We Smart Enough to Know How Smart Animals Are? convinced Xie to drop human arrogance: intelligence evolves continuously, animals use tools, recognize themselves in mirrors, play politics like chimpanzees in de Waal’s work, and reason about others’ minds. In one experiment a chimp that sees a researcher eating a banana goes straight to the box with the apple.
The target is still human-like intelligence, because it benefits the world most. But he is inspired by Sutton’s retort to LLM triumphalism: building a squirrel’s intelligence—goals, intrinsic reward, hunger, emotion, social life, survival in the real world—is the truly hard problem, and once you can do that, writing code or going to the moon is easy. A 12-year-old can do nearly every household chore; no robot can. That needs the “second half of pre-training,” which robot startups can’t afford because they pour money into hardware scaling. Its inputs will be continuous, noisy signals—starting with video—while the right training objective remains an open research question.
People matter enormously to him; the podcast spends hours on the researchers who shaped him, from Hou Xiaodi and Tu Zhuowen to Kaiming He, Fei-Fei Li and LeCun. He calls how they came together a law of attraction—streams converging into one river. His motto, borrowed from Liverpool manager Jürgen Klopp answering José Mourinho’s “I am the special one,” is “I’m the normal one.” He wants to be the team’s battery, as LeCun is for him.
He says he feels frustrated every day—research means groping in the dark, with maybe 5–10% of the time spent in the joy of results—but this AI era, with so much open discussion, makes the search less lonely. LeCun is relentlessly optimistic, perhaps because he lived through the AI winters and was proved right; he often says a small group can always see where technology is going while most people are busy elsewhere.
Xie urges researchers to love real life and avoid an echo chamber. In New York he walks through Washington Square Park—pianists, dancers, chess players, students—and is reminded the world is bigger than AI, which raises the question of researchers’ responsibility. His recommendations include the series Person of Interest, Pantheon (based on work by his fellow townsman Ken Liu), the AI short film The Total Pixel Space, and two books: Gödel, Escher, Bach and Zen and the Art of Motorcycle Maintenance. What matters most, he says, is sincere communication between people; research is a form of it, and trust built through papers even helped AMI’s fundraising, when a Black Forest Labs founder he had met once urged an investor to back him. Can we predict fate if the world is a world model? No—we would need a computer the size of the universe, and the answer might be 42.
NVIDIA’s Ming-Yu Liu, asked about this interview, disagrees that the Valley is simply hypnotized: the leading companies all started from language models, and coding agents are too useful to dismiss.
Read: “You don’t need to beat every rival”: NVIDIA’s Ming-Yu Liu on Cosmos and making markets →Listen to the full conversation
Episode 133 · about 6 h 45 min, in Mandarin, by Zhang Xiaojun (张小珺). Our pages are an English retelling and analysis, not a word-for-word translation.