“You don’t need to beat every rival”: NVIDIA’s Ming-Yu Liu on Cosmos and making markets
The first NVIDIA researcher promoted to vice president leads Cosmos, its world foundation model. Jensen Huang’s instruction: keep going “to Cosmos 97.”
In 60 seconds
- Jensen Huang’s rule, in Liu’s telling: build “zero-billion-dollar businesses”—markets that do not exist yet—rather than fight over existing ones. Physical AI is that market.
- Cosmos started around March 2024, after Sora. Its goal is not content creation but robots and self-driving: better data, a better starting point, and better training environments—“a Matrix for robots.”
- Cosmos 3 folds four earlier models into one that takes language, video, audio and actions. The model, training framework and some data are open.
- Liu reads OpenAI’s retreat from Sora as a focus decision: Anthropic’s coding agents showed where the token economics are.
- His prediction: “A ChatGPT moment for physical AI is coming.”
Ming-Yu Liu (刘洺堉), NVIDIA’s vice president of research and the first NVIDIA researcher promoted to VP. Born in Taiwan, 20 years in the US and 10 at NVIDIA, he leads Cosmos, NVIDIA’s world foundation model for physical AI—about 80–90 people report to him and 200–300 work on it. He has practiced Chinese martial arts since 18.
NVIDIA sells the compute everyone else trains on. How its own research leader thinks about competition, open models and robots tells you where the platform wants the market to go.
Five ideas worth your time
Liu frames model competition as not a “Squid Game”: there are too many problems to solve. He expects model capabilities to converge, with differentiation coming from ecosystems, the way software companies differentiated once everyone could write code. CUDA, he says, created today’s AI market; the question is how to create the next one. When the interviewer notes rivals will say NVIDIA just wants to sell GPUs, he laughs: you need compute anyway.
Robots learn by interacting with the world, which is slow and expensive. A good world model can generate the data you could not collect, give a robot policy a better initialization, and—the hardest part—provide many realistic environments to learn in at once.
That became Cosmos 3: one model that merges reasoning, prediction, transfer and robot policy, taking language, video, audio and action as inputs (two “towers” for discrete and continuous signals, for now). He cites the Art of War—spread your forces and each is weaker—and the same logic for why OpenAI pulled back from Sora to focus on coding.
NVIDIA open-sourced the model, training framework and part of the data because many frontier labs now publish little. He credits DeepSeek’s fast iteration as an influence and praises Chinese labs for hiring many interns, which spreads know-how. He sees China’s manufacturing base as an advantage for robots that integrate hardware and software.
He describes NVIDIA’s top-five priority emails, which Jensen reads to spot signals; teams aligning around a mission rather than an org chart; and “torturing people to greatness.” Jensen once asked him, after he blamed an unfair benchmark setup, “Are you a quiet baby?” He stopped making excuses.
“You don’t need to beat your rivals.”— Ming-Yu Liu
“Are you a quiet baby?”— Jensen Huang, as recalled by Liu
What it means for you
- Don’t train a world model from scratch if an open foundation model gets you 80% there. Spend on the data and deployment only you have.
- Liu’s “zero-billion-dollar” test is useful: if someone already owns the market, it may not be worth entering.
- Coding agents’ token volume reshaped OpenAI’s priorities. Price your product around long-running agent work, not chat.
- NVIDIA open-sourcing Cosmos commoditizes one layer of physical AI. Value shifts to proprietary data, hardware integration and deployment.
- China’s manufacturing depth is a structural edge for robot companies that integrate hardware and software.
- Liu’s bet—a “ChatGPT moment” for physical AI—would mean a large step-up in compute demand from robots and cars.
- Physical AI is, in his words, a good time for researchers to enter; so are materials and drug discovery.
- Cosmos is built for post-training on your own video-action data—check whether that fits your robot or pipeline.
- Evaluation that matters in physical AI is with experts who have the problem, not only benchmarks.
Read it chapter by chapter
Every argument, example and number from the original, retold in our words in the original order. Short quotes are attributed.
- 01 · On competition: “I have no plan to kill anyone”
- 02 · On boundaries: “the ability to be number one—used for the greatest contribution”
- 03 · On Cosmos: “Just do Cosmos 97”
- 04 · On Cosmos 3: “Louis Vuitton sold more after showing fewer products”
- 05 · On “The Matrix”: pain and suffering drive progress
- 06 · On Jensen Huang: “30 days of cash left”
- 07 · On himself: “not to show destructive power”
Liu rejects the idea that model competition is a Squid Game. If everyone tried to solve exactly the same problem, it would be; but there are countless problems and tastes. Should every company train its own model? Treat it like electricity: early factories built power plants to guarantee supply, but once power is abundant you don’t. And people don’t stay aligned forever—he notes that Dario Amodei left OpenAI and asks whether everyone at Anthropic believes the same thing. Anyone with conviction and resources can build.
His controversial forecast: model capabilities will converge, and differentiation will come from what companies build around them. Leaders have more time to find that edge, and strong companies can widen their ecosystems. Software was once rare too; the survivors found other ways to stand out. He likens deep learning to traditional Chinese medicine—a black box refined by many trials into a body of knowledge that spreads.
Is world modeling a war NVIDIA can’t lose? Not if others also open-source models, he says. Physical AI is a huge market: if every home had two or three robots, each needing compute, demand for compute would soar, which is enormously good for NVIDIA. The goal is to make the market, as CUDA made today’s AI market. He quotes Jensen Huang’s idea of a “zero-billion-dollar business”—no business today, a billion-dollar one if it succeeds; if someone already does it, maybe don’t. A $100-million-a-year business would be a distraction for NVIDIA. That is risky: CUDA meant selling programmable GPUs at a premium, but without it GPUs become a commodity.
Liu believes he is the first NVIDIA researcher promoted to vice president. He is VP of research and of the Cosmos lab; 80–90 people report to him directly, and 200–300 work on Cosmos in total through dotted lines, with him as the mission owner. He still reads papers, reviews code and argues details—he fears losing touch with the problems as a manager.
Researchers who later built Sora once interned on his team. He calls that a source of pride, not regret, while asking himself what he failed to see at the time. Why did OpenAI largely shut Sora down? He frames it as focus: Sam Altman runs a portfolio like an investor, and over the past year Anthropic’s coding agents showed enormous economic value—chat produces a minute of tokens, coding agents work for 24 or 48 hours. OpenAI concentrated on coding. Sunzi’s warning applies: prepare everywhere and you are weak everywhere.
Why does an infrastructure company do research? To build infrastructure applications actually need, you must walk their path; chips and system architectures take years to plan. He also wants ambitious young people everywhere to have tools, which is why NVIDIA open-sources models—mostly free, with a long road left to real deployment, especially in embodied AI. Shipping generation after generation signals commitment so partners can invest in their own hard problems. On whether he would rather build the number-one model or avoid competing with customers: “number one” is fleeting. He wants the capability to build a frontier model, used for the greatest social contribution—not a model for a few people that concentrates power.
His group has a single project: Cosmos. After version 1 he asked Huang whether to continue; Huang said to go all the way to Cosmos 97. The number was offhand, but the commitment was real.
Why world models for physical AI, not LLMs or Sora-style creative video? His background is computer vision and media; NVIDIA has an LLM team (Nemotron). Sora was a world simulator aimed at creators, but many social problems need physical tools: self-driving cars that keep older people mobile, robots for dangerous or labor-intensive work, help with chores. So Cosmos went toward physical AI.
His homepage vision—“a Matrix for robots”—is about generalization. Cosmos helps in three ways: better data (synthetic data covering what the real world didn’t capture), a better starting point (a world model’s representation bootstraps robot policies), and better environments (many realistic worlds to practice in at once, the hardest part). The project was greenlit around March 2024, after Sora made the case obvious to Huang, who is disciplined about pacing. First-mover advantage depends on your position; the most profitable video-generation business today, he notes, is probably ByteDance’s Seedance.
Huang insisted Liu propose a name himself because naming defines positioning: Cosmos, a universe, aimed at solving physical problems rather than content creation, and at developers rather than consumers. Physical AI is a necessity: everyone ages, labor is short everywhere. His definition of world models is practical: models built for prediction, for understanding why things happen (like diagnosing a stalled production line), or for 3D reconstruction. He no longer works on reconstruction. The term is too broad to define precisely, which he doesn’t mind. Cosmos was named a “world foundation model” so others can build their own world models on it. NVIDIA also invests in LeCun’s AMI and Fei-Fei Li’s World Labs; he sees NVIDIA as an enabler, and its real competitors as companies building an entire AI ecosystem.
Cosmos moved fast, inspired by DeepSeek’s three iterations in a year. Working with partners, the team built models for prediction, for making simulator footage more realistic (Transfer), for explaining events (Reason), and then a robot policy model, after realizing that how pixels change is close to how actions change. Because all of these model the same world, a foundation model should absorb all their data—so they moved toward a single model.
He once told Huang he would need about 22 models. Huang praised the effort, then told him Louis Vuitton’s sales rose after it cut the number of products on display: too many choices confuse customers. Liu took it to heart; each model must be maintained, and spreading effort slows iteration. Cosmos 3, started around October last year after Cosmos 2 and 2.5, merges language, video, audio and action in one model. A Chinese founder said the paper had few big new ideas but verified many engineering details; Liu largely agrees and says details are what make models better—Mixture-of-Transformers and related ideas existed, but scaling them to combine audio, video and action is the hard part.
NVIDIA released the model, training framework and some data, because many frontier labs no longer explain key steps. Action is a first-class citizen because physical agents change the world. Modalities help each other: predicting video helps action learning, audio marks contact moments, detailed language helps video. Two towers handle discrete signals (better for understanding) and continuous ones (better for generation), mainly so customers can post-train one side without degrading the other; a single tower is the next goal, but usability comes first. Would it beat Claude at coding? Again: spread yourself thin and you are weak everywhere—strategy is deciding what to give up.
Evaluation needs benchmarks, arena-style comparisons and, above all in physical AI, experts with real problems. Data is split into navigation (easier to share across robots) and manipulation (harder); the trend toward human-like hands led them to add first-person video of human hands to Cosmos 3. NVIDIA’s edges: cheaper compute, years in physical AI (Thor chips, Isaac simulation), and a genuine interest in customers’ success. Compute for Cosmos is in the tens of thousands of GPUs. He is proud Cosmos 3 beat all its predecessors; the regret is shipping before it fully converged, which 3.1 and 3.2 will address.
The Matrix metaphor is about learning faster in simulated worlds, not about machines enslaving people. Society is built and operated by people, and with consensus, catastrophe can be avoided. He doubts chips feel the pain and suffering that drive human progress, while supporting guardrails.
To rival world-model teams: world models are tools; work with NVIDIA and go further and faster. Yes, NVIDIA wants to sell you GPUs—you need compute anyway, and cloud companies bought lots and made more. Could Cosmos be the next CUDA? CUDA turned application logic into GPU computation; future software is AI built on AI, so Nemotron and Cosmos could be part of it—but only as part of a larger stack. CUDA’s path, he thinks, can’t simply be copied.
On China’s models: very impressive. U.S. frontier labs rarely take interns now, keeping know-how inside; Chinese companies take many, spreading knowledge. On Xie’s “hypnotized” claim: he disagrees—today’s leading companies all started from language models, and he can’t imagine working without coding agents. He expects 2026 to be a year of upheaval: SpaceX’s IPO, OpenAI and Anthropic heading toward listings, soaring San Francisco housing. His prediction: a ChatGPT moment for physical AI is coming. On robots: in the U.S., only Tesla and Figure build full robot bodies; China’s manufacturing allows tighter hardware–software integration, which is worth watching.
In his ten years, NVIDIA went from unknown to his parents to synonymous with AI. There were hard times—when Google’s TPU arrived and the stock fell, some colleagues left; NVIDIA kept beating TPU on MLPerf until it was the only one still entering. Huang’s hair has turned white, but he still reads employees’ top-five priority emails for signals. The company has more than doubled; GPUs moved from desktops to data centers.
Huang says, every time, that the company has only 30 days of cash left—not literally, but to force the right decisions. During the pandemic downturn, NVIDIA cut travel and tools (even considering dropping Slack or a cloud provider) rather than people. Liu hasn’t seen a layoff in his decade; there is no forced ranking, only performance reviews. The downside is slower transitions when technology shifts; the upside is trust and cooperation. NVIDIA runs marathons—CUDA took ten years—and doesn’t pit teams against each other; it trusts people with a mission and sets a very high bar, in Huang’s words torturing people to greatness.
NVIDIA’s distinctive practices, as he lists them: no layoffs, top-five emails that surface signals and duplicate work, radical transparency (Apple is the counterexample), and “the mission is the boss,” with teams aligning like an organism. Huang shaped him: he sent papers for Liu to summarize, learned computer graphics and then deep learning from scratch, decided from first principles, and prioritized ruthlessly. Liu now switches topics every half hour. His computer-science lesson: be lazy—don’t compute what you don’t need. Once, when Liu blamed an unfair experimental setup for losing to a rival, Huang asked, “Are you a quiet baby?” There are no excuses; get the conditions you need.
Is his low ego real or a mask? Confidence, he says—the belief that what they build helps others. As a young researcher, trained to prove state-of-the-art results, he was highly competitive; now he knows every state of the art is soon surpassed, and he cares about developing people and hearing that Cosmos helped a partner. You don’t need to beat your rivals; he prefers to call them partners.
He wakes around 4–5 a.m. and sleeps around 10–11. He has practiced Chinese martial arts since 18 under Xu Ji, in a lineage traced to Li Shuwen, known for Bajiquan and Piguazhang, and recently took up ice plunges. Martial arts, inspired by Jin Yong’s novels, taught him control: a punch’s power runs from the feet to the hand in an instant, which demands constant self-correction—“not to show destructive power.” Favorites: beef noodle soup and the Bay Area’s weather. His advice to young researchers: stay calm and confident, don’t fear that AI will end next year, and look at physical AI, materials and drug discovery. How should history remember him? Don’t talk about him—he hopes Cosmos helps a few important physical-AI builders reach their ChatGPT moment.
Saining Xie argues the opposite about language: that it is a crutch that is already polluting vision. Liu’s answer is pragmatic—use language to talk to people, video and audio to learn physics, and actions to change the world.
Read: “Silicon Valley is LLM-pilled”: Saining Xie on why world models need the world →Listen to the full conversation
Episode 150 · about 3 h 36 min, in Mandarin, by Zhang Xiaojun (张小珺). Our pages are an English retelling and analysis, not a word-for-word translation.