Technology
Paragraphs

OVERVIEW
 

The global demand for “AI sovereignty” is increasingly shaping AI policy and safety discussions. What was once a niche concern has become a widely shared priority, with more countries seeking control over how AI is built, deployed, and governed within their borders. This memo synthesizes key insights on the drivers of AI sovereignty, how countries are operationalizing it in practice, and the implications for U.S. diffusion strategy and managing global risks.

Three premises frame the analysis. First, emerging economies and middle powers are becoming crucial partners, consumers, and suppliers in the global AI ecosystem. Second, U.S. leadership in frontier AI creates a narrow—and likely diminishing—window to shape how this technology diffuses. Third, understanding what countries actually want from AI is critical to designing a sustainable U.S. diffusion strategy and navigating shared risks.

These insights draw on ongoing research from the Carnegie Endowment for International Peace (CEIP) and Stanford's Program on Geopolitics, Technology, and Governance (GTG), and were further developed at a workshop, “AI Sovereignty, Diffusion, and Risk: Strategy in a Contested Landscape,” convened by CEIP and GTG on June 17, 2026 with experts from civil society, industry, and academia.
 

KEY INSIGHTS
 

“AI sovereignty” is a politically potent but analytically ambiguous concept, driven primarily by a desire to manage dependence and extractivism, and to obtain AI suited to local contexts.
 

While the term “AI sovereignty” is ambiguous, it has become an increasingly influential part of international AI discussions. Sovereignty goals seem less about achieving true technological autarky (widely seen as unfeasible), and more about a desire to secure national interests (economic, security, cultural) within a tech landscape dominated by the U.S. and China. In this way, AI sovereignty might be practically understood better as “AI agency:” a country's ability to execute choices in line with national interests. In search of this agency, countries will likely pursue a strategy that combines assured access arrangements for foreign AI models and hardware with efforts to diversify and build domestic capacity (indigenous or modified foreign open-source) if such access is disrupted. These domestic capacity efforts tend to focus on cheaper, multilingual, and multimodal models contextually attuned to underserved markets.

Sovereignty can serve as a source of resilience, and, increasingly, may serve as a deterrent against being cut off from frontier AI capabilities. To achieve this deterrent effect, countries are likely to seek sources of leverage along the AI value chain. These may include:

  • Control over scarce resources in the AI supply chain, such as critical minerals or low-cost energy for powering data centers.
  • Significant market power that makes a country an indispensable consumer of AI services.
  • Refining high-value local data for domestic value capture.
     

These sources of leverage, however, could be a depreciating asset, as a nation may only be able to threaten or use its leverage once before providers diversify away from it.

Perceptions of AI risk within sovereignty debates rarely focus on concerns over shared catastrophic and large-scale threats.
 

Discussions of AI sovereignty in many middle powers and emerging economies are primarily driven by concerns over dependency, extractivism, market concentration, and geopolitical subordination. Cross-border, large-scale AI risks like cyberattacks, biosecurity, or rogue AI agents have not been central to these debates to date, as they are often perceived as remote or lower priority than immediate economic and political concerns and assured access to AI technology.

Recent events, such as the U.S. government blocking Anthropic's Fable model access, have reinforced this focus on dependency. While the incident highlighted the potential for dangerous capabilities, the primary international reaction has been concerned with the U.S. wielding a “kill switch” over model deployment, reinforcing fears of unilateral control and strengthening the case for AI sovereignty.

This dynamic could shift as AI capabilities diffuse. The widespread availability of powerful models capable of causing significant cross-border harm, such as cybersecurity failures in critical infrastructure, may force a re-evaluation. In such a scenario, the salience of shared safety and security could rise, potentially aligning national interests more closely with global risk mitigation efforts.

Countries are actively pursuing sovereign AI projects, but a significant gap persists between ambitious strategies and operational capacity.
 

A clear trend of sovereign-related AI projects and announcements is underway globally, especially in the EU, Indo-Pacific, and Gulf States. These initiatives range from building national compute clusters and developing domestic models to pursuing legal arrangements to ensure data localization and provide assured access to foreign AI models.

However, a significant gap often exists between high-level ambitions and the operational capacity to implement them, as many efforts lack clear funding and technical expertise.

The viability of these national ambitions may hinge on a critical, and often unresolved, distinction: identifying which use cases can use “good enough” AI, relying on less advanced models and infrastructure, versus those that require access to the frontier.

The future trajectory of AI—whether dominated by a few frontier labs or a broader open-source ecosystem—is a central uncertainty shaping national strategies.
 

A fundamental tension exists between two potential AI futures: one where a handful of frontier labs create a runaway capability gap, and another where open-source models remain competitive and useful for most purposes.

In a world where frontier models pull away, a U.S.-led stack could become central to the global economy and international security, giving the U.S. unprecedented insight and international leverage.

In contrast, in a world where open-source and fast follower models remain competitive, the strategic calculus shifts, potentially lowering the stakes of frontier competition for many countries and enabling more diversified technology ecosystems.

The concept of an organized “third stack” as an alternative to U.S. and Chinese ecosystems built on open-source models and diverse hardware is a potential pathway for middle powers seeking to avoid dependency. This alternative is only viable if a coalition of middle powers align on standards for procuring and deploying AI to pool their collective purchasing power; the European Union’s regulatory and tech sovereignty efforts are informed by this logic. However, multi-state organization around a third stack faces significant collective action problems, and any effort to create an independent stack could be viewed as a challenge to U.S. national security, incentivizing Washington to pursue bilateral engagements that thwart an emergent alternative.

Even with open models, countries may still want assured access to frontier AI for limited, exquisite capabilities.
 

Even if many countries' economic, development, and security goals can be achieved through “good enough” AI, countries may want access to high-end capabilities for specific purposes. This could give the United States an opportunity to offer assured access to frontier AI in exchange for safety and security commitments.

However, an assured access framework faces major challenges:

  • Trust in U.S. assurances: The U.S. is often not currently seen as a credible, long-term partner. Without trust, assurances are secondary to mutual leverage.
  • Third country leverage: A security–frontier bargain appears most viable with countries that possess something the U.S. wants (e.g., India's data, Brazil's energy resources). Leverage is unclear for states that lack such bargaining chips.
  • Risk perception gap: Safety commitments, especially those framed around catastrophic risks, do not resonate strongly in many parts of the world, where developmental and economic priorities are paramount. For Global Majority economies, legible near-term AI risks include displacement of labor within domestic informal sectors and the devaluation of labor within global value chains.
     

Significant doubts about the U.S. government's ability to execute a nuanced AI partnership strategy highlight the roles of market forces and non-state actors.
 

Many believe the U.S. government currently lacks the capacity and planning to execute a complex global AI diffusion strategy. The U.S. government’s existing toolkit for promoting the U.S. tech stack, including bodies like the Development Finance Corporation (DFC) and EXIM Bank, is difficult to deploy effectively. The government's most effective role may be to de-risk investment and lend credibility in geopolitically critical areas where a natural market does not exist.

In the absence of a robust government strategy, market forces are the primary driver. The quality of U.S. technology creates a natural pull, but this may be counteracted by U.S. policy unpredictability.

In this vacuum, non-governmental actors are stepping in. Philanthropic organizations are brokering agreements between U.S. AI labs and Global Majority countries for specific use cases. However, it remains unclear if these ad-hoc efforts can scale into a coherent ecosystem that benefits U.S. interests.

PRIORITIES FOR FURTHER RESEARCH
 

This analysis surfaces several unresolved questions that warrant further research and debate. Priorities for future inquiry include:
 

  • Clarifying the competing futures of AI: Under what conditions might frontier AI models achieve a decisive, compounding advantage over open-source alternatives? What are the key technical and economic indicators that policymakers should monitor to assess which future is becoming more likely? What are the implications for AI risk management?
  • Mapping points of national leverage: Beyond theoretical control over chokepoints, what forms of economic, political, or geographic leverage have proven most effective for middle powers in securing favorable terms for AI access? How durable is this leverage, and can it be pooled regionally to overcome collective action problems?
  • Designing a viable U.S. partnership model: What specific, actionable policy tools are required for the U.S. to offer a compelling “assured access” bargain? How can such a policy be designed to be credible across administrations, especially for countries that lack significant intrinsic market or resource leverage? How should AI risks factor in?
  • Integrating safety into diffusion frameworks: How do perceptions of AI risk influence national sovereignty strategies? What practical mechanisms can embed safety and security commitments into technology partnerships?
     

Download the full PDF here.

All Publications button
0
Publication Type
Policy Briefs
Publication Date
Journal Publisher
GTG–CEIP
Authors
Melissa Morgan
News Type
Commentary
Date
Paragraphs

A breach of Hugging Face's servers. AI safety debates on Capitol Hill. New models like Kimi K3 coming online. There's been no shortage of headline grabbing AI news and debate in the last few weeks. To talk through these topics and how they relate to American national security, the heads of national security policy at OpenAI and Anthropic joined Colin Kahl on the World Class podcast.

Together, Sasha Baker (OpenAI) and Tarun Chabra (Anthropic) discuss how artificial intelligence is intersecting with security and defense strategies, the U.S.-China AI race, and how the risks and advantages of AI are starting to reshape national security.

Sasha Baker is the head of national security policy at OpenAI. Prior to that, she served as the acting under-secretary for policy and the deputy under-secretary of defense for policy at the Pentagon, and as the senior director for strategic planning on President Biden's National Security Council staff.

Tarun Chhabra is the head of national security policy at Anthropic. He previously served as deputy assistant to the president and coordinator for technology and national security on the National Security Council staff, where he coordinated the Biden administration's strategies for technology competition with China and technology partnerships with U.S. allies and partners. 

This episode's reading/watching recommendations are AI 2027, a project by Daniel Kokotajlo; AlphaGo, a documentary by director Greg Kohs; and "Nineteenth-Century Horse Sense" by Francis Thompson.

TRANSCRIPT:


Kahl: You're listening to World Class from the Freeman Spogli Institute for International Studies at Stanford University. I'm your host, Colin Kahl, the director of FSI.

I'm very excited to welcome my good friends and former colleagues Sasha Baker from OpenAI and Tarun Chhabra from Anthropic for what promises to be an insightful conversation on the good, the bad, and the ugly of how artificial intelligence is intersecting with national security. We're going to discuss the U.S.-China AI race, the risks AI poses to security, and the opportunities and advantages AI might generate in the national security space.

These are fantastic guests to help us grapple with these complex topics. Sasha Baker is the head of national security policy at OpenAI. Prior to that, she served as the acting under-secretary for policy and the deputy under-secretary of defense for policy at the Pentagon, and as the Senior Director for Strategic Planning on President Biden's National Security Council staff.

Tarun Chhabra is the head of national security policy at Anthropic. He previously served as deputy assistant to the president and coordinator for technology and national security on the National Security Council staff, where he coordinated the Biden administration's strategies for technology competition with China and technology partnerships with U.S. allies and partners. 

Sasha, Tarun, thanks for coming on to World Class. 

So look, I just told everybody your job titles; they’re very fancy. People know your companies. But perhaps we could start with describing what your roles at OpenAI and Anthropic actually entail.

Sasha, maybe let's start with you. What do you actually do on a daily basis?

Baker: Well first of all Colin, thanks for having me. It's fun to be here with two former colleagues. And what you didn't mention in my bio is that in my deputy role I was the chief “Colin Minder” of the Pentagon for a number of months.

What do I do? It's the most interesting job maybe I’ve ever had. I get to talk to governments all around the world every day about basically two things. The first is: what is the opportunity space look like as it relates to AI being used in the national security domain? So, how can we help governments that want to use these tools to make their populations safer? How do we help them do that?

And then the second thing that I talk to governments about is the AI risk space and the ways in which AI could up-level or enable a bad actor to do something that we really wouldn't want to see. So, I spent most of my career in government, as I think Tarun did as well. And what's fun about this job is it allows me to still be part of the same kinds of conversations that you and I would have had when we were serving at the Pentagon, but just from a totally different vantage point.

Kahl: Tarun, how about you?

Chhabra: Let me add my thanks, Colin, for having me too, and it's great to be here with Sasha and with you.

Obviously a lot of similarities in my role as well. I think of it as trying to help prepare policymakers for what we see as coming down the pike in AI development: to help them prepare and whether that's folks who are doing reporting and analysis, or whether that's folks who are making policy, or folks in Congress. We want them to know what's coming so that they're ready.

And then the second part, of course, is making sure we can get the best possible technology into their hands for national security purposes. As Sasha said, that's first and foremost with the U.S. government, but it's also, of course, with our closest allies as well.

Kahl: Great. Let's jump right into the types of conversations you're having. I'm sure as you're talking to officials around the world, there's a lot of focus on the U.S.-China AI competition. I think as many of our listeners will know, in recent weeks, Chinese AI labs have released some very impressive models, including ZAI's GLM 5.2 a few weeks back, and most recently Moonshot's Kimi K3. Obviously, people are familiar with very good models released by other Chinese labs like Deep Seek.

Tarun, maybe starting with you on this one: how would you assess the current gap between the best models coming out of Chinese AI labs and the frontier AI models coming out of labs like Anthropic and OpenAI? Is the U.S. ahead? If so, by how much? How would you assess the race right now?

Chhabra: I think our view's been pretty consistent even as new Chinese models have come out, which is, we believe that we remain six to nine months ahead of Chinese models at the frontier. And we think that really matters. So if you think about having really advanced cyber capabilities, having six to nine months with those superior capabilities really matters from a national security perspective.

That being said, we think that six to nine months owes a lot to the fact that leading Chinese AI model developers are distilling our models, those of U.S. frontier companies. And without that distillation, we could probably have a lead closer to around 18 months. And that would be even better, obviously, from a national security standpoint as well.

So that's why we've really appreciated the work that Sasha and colleagues at OpenAI have done to expose some of the distillation that is happening, why it matters, and we've really appreciated recent steps and pronouncements by the current administration to say this really is a national security issue that we need to be taking taking seriously.

It is important to note that distillation doesn't take you right up to the frontier. It keeps you behind. But the time really matters, and distillation is having a big impact.

And I think one reason why we see more attention to it right now is that the absolute capability you get from distillation matters too. And so as models become more and more advanced and more capable, even if we maintain that six to nine month gap, once you hit certain thresholds of absolute capability, that does become a concern.

Kahl: So just so that our listeners are following along. When you say distillation, really what we're talking about is a Chinese lab basically pre-trains their model. They run a big training run, but then as they're post training and fine tuning their model, they're actually engaging in a lot of queries back to say Claude or some version of GPT and using the answers from that to essentially reinforce the learning of their model and fine tune it.

Is that a fair description of distillation?

Chhabra: That's right. And as you have more capability baked into the model at that later stage of model development, the more opportunity there is for it.

I think it's important to note, however, that compute still matters. You can steal the recipe, but you still need a kitchen. So this is why we've been very vocal on the need to maintain export controls, whether it's on manufacture chips or the chips themselves, because that's a limiter. And many folks are still accessing these models for inference through APIs.

And in terms of the business model, even when the models are open, often these same companies are using their compute as a way to fund what they're doing.

It's really important to say this is industrial espionage. We've seen the playbook before where you have heavy, heavy subsidies from the Chinese state going into various industries to try to scoop up as much market share and then hold on to it for as long as possible. Except here I think it's not just a national competitiveness and economic issue, there are real national security consequences too.

Kahl: I want to come back to the export control issue in a second, but Sasha: does OpenAI generally share the assessment that the best models being produced by your company, by Anthropic, by Google Deep Mind are six to nine months ahead at the frontier? And do you share Tarun's concern about China basically being able to be a fast follower in part due to distillation?

Baker: Yeah, I think somewhere around the six month mark in terms of the lead is probably my best guess. Maybe there are just a couple points to make in addition to what you heard from Tarun.

The first is that not all distillation is bad. We distill our own models to create fine-tuned or fit-for-purpose versions of a model. And we allow developers to do some forms of distillation on our platform.

What we're really concerned about is what we would call ‘adversarial distillation’, which is unauthorized attempts to extract the capabilities of a U.S. frontier model in order to build something that then is kind of a competing system. And that is what I think has economic national security concerns.

I do want to be clear: there are a lot of very, very talented AI researchers in China, and I think it is the case that they would have very capable models even absent distillation. But this certainly does give them a leg up.

And I think the real thing to be concerned about here is actually the safety stack that comes from those distilled models. Because oftentimes they will distill a capability, but they won't export the safety stack. And so when you look at some of the models coming out of China—and this is true whether they're open or closed—you find that they are highly permissive in allowing for tasks that U.S. frontier labs spend tremendous amounts of energy trying to prevent.

And that has implications for the overall threat picture for global governance. And it's something that we oftentimes talk to them about. So that's an area where I think that there are real opportunities for us to do more together.

Because what all three frontier labs — so, Google, OpenAI, and Anthropic — have all put out assessments of where we see distillation happening on our platforms, but we can only see what's happening on our platform. And it requires that partnership with government and partnership with each other to be able to get a sense of the fuller ecosystem around this.

Kahl: All very interesting. Both you and Tarun have talked about distillation. My sense is that the fact that China is able to only be six or nine months behind . . . maybe that's partly due to distillation. 

I think there's also a view that they do have very smart engineers who have engaged in innovation in terms of algorithms. But they've also used smuggled NVIDIA chips, very powerful Blackwell chips have been used to train some of these models in illicit data centers. You also see reports of using, essentially, remote compute access from data centers in places like Malaysia to train these models.

I think a lot of that comes back to this debate about export controls, right? The first Trump administration put in export controls—very important ones—on semiconductor manufacturing equipment, especially advanced lithography equipment that are necessary to produce sub seven nanometer chips.

The Biden administration then layered on a lot more export controls on semiconductor manufacturing equipment and tools, and then also put controls on the sale of advanced chips directly to China. And yet China is still able to kind of fast follow.

So I guess the question is: does that suggest that export controls ultimately are always going to be imperfect? Maybe they're a fool's errand? They're not worth the cost? Or does it just suggest this is a game that you have to keep playing and there are ways in which the export controls need to be tightened.

And we'll start with you, Tarun. You've thought about this more than just about anybody I know.

Chhabra: I think the way to think about this is as a counterfactual. What if there had been no controls in place? Where would we be? And to the point you just made and that Sasha made earlier, China has tremendous AI talent and as you know also, they have a tremendous amount of energy that's coming online, it's something like 7 to 8x what is coming online in the United States, and that's before you get to the nuclear build out. 

But the one problem they have—and don't take it from me, take it from the leaders of China's top AI labs—is compute. That matters both for model development, but also for serving the models in terms of the race to to eat up global market share. 

And even with the latest releases over the last two weeks with GLM 5.2. and Kimi, you see already a problem in serving the level of demand to date. And again, one lab leader after another from China complains about the access to compute.

So absent the controls, could we be in a situation where China's in the lead? I think it's very possible.

Kahl: I think it's an important distinction you draw for our listeners who maybe aren't quite as in the weeds. Obviously there's all the computing resources: the data center is full of tens of thousands or hundreds of thousands or maybe even millions of leading edge AI accelerators.

But you also need compute to serve those models, that is to run inference. So anytime you pull out your smartphone and you prompt Claude or ChatGPT or Gemini, it's going back to a data center somewhere to run that query. And the more advanced the models, the more compute they require for inference to run really complex tasks. So compute obviously matters there, too.

I wonder, Sasha . . . you have all have discussed distillation as essentially IP theft. You also mentioned that these distilled models may not have the safety guardrails that some of the models that you all are producing. And we know that those are imperfect as they are.

Do you get a sense that the U.S. government is trending towards thinking about regulating Chinese models in some way? That is, either putting Chinese companies on an entity list or telling U.S. hyperscalers they can't serve Chinese models? Do you get a sense that the administration is thinking about clamping down on Chinese models because of so many of these issues?

Baker: I'm not sure we know any more about the answer to that question than you might also read in the newspaper.

What I can tell you about the conversations that we have with the U.S. government is right now we are talking with them about how do we create a mechanism of evaluating models—not just Chinese models, but American models or you know, models from around the world—so that we have a collective and common understanding of what we're even talking about here.

What are the capabilities of these models and how do you measure them? What are the safeguards around these models? How do you measure that? How do you determine what is sufficient? 

And I think that there's a really important role that the U.S. can play and U.S. leadership can play globally in helping to define some of those questions and create processes that will allow governments around the world to understand the landscape a little bit better. And then each government, I think, is going to make its own determinations about what they want to do with that information.

Kahl: One of the things that distinguishes a lot of the leading Chinese models is that many of them are open source or more precisely open weight in the sense that, you know, they can be downloaded and their parameters can be further modified or fine-tuned by users on their own servers.

And even when these models are accessed directly, at least from what I read, it seems like their API costs tend to be pretty low and their token usage tends to be very efficient.

In contrast, it seems like the best U.S. models tend to be closed weight, they're proprietary models, although there are some good open weight models that are being released by companies like NVIDIA and Thinking Machines.

But I guess the question I have, maybe Sasha starting with you is: what's the business model here for Chinese firms? They still have to spend money to train these models. They have to buy compute or lease compute to do it. They have to pay the salaries of all these really smart people who are working in these labs. And then they are essentially giving away their technologies for free or at very low pricing. How are they going to stay in business?

Baker: It's a super good question. Before I try to answer it, let me just say up front: we have always thought that there's an important role for open source models in the AI ecosystem. We have one ourselves. We think that they play a really important democratizing and innovation role. And there's room for open source and there's room for closed source models.

But having said that, there is the question about like, well, how do you actually make money off of a model if you're giving it away? And I think Tarun hinted at that answer earlier, which is that the model, in some cases, is actually not the product, right?

So if you think about some of these large Chinese labs — think of an Alibaba, for example — the model may be a loss leader that incentivizes other businesses into the Alibaba ecosystem, whether that's the cloud or what have you.

And then for some of the smaller labs, to Tarun's point, they may not charge for the model, but they can charge for the convenience, right? If you want to use the API, if you want access to their harnesses, etc. And there's a stickiness there that then gets people to kind of come back again and again.

At a nation-state level, I do think there's an element here, which is that the open weight approach is not just about having a standalone business model, it's also a distribution strategy that allows for the capture of market share. Because as I said, there's some stickiness to this. So we think competition is good; we compete, of course, across the American labs, we compete internationally. And that's healthy; it drives innovation.

We expect that businesses will evaluate and use a wide range of models. And we feel pretty good as a company about our value proposition, which is we have a model that is secure, that is reliable, that can deliver at scale, and that generates what we think is more useful work per dollar per token on a more reliable level than others that are out there.

And we feel good about that. And our customers tell us that that's something that sets OpenAI apart. So we're going to continue to try to do what we think we do best.

Kahl: Tarun, Sasha mentioned the state level, and you've thought a lot about what Beijing is trying to accomplish at the nation state level.

In one sense, open weight models are a good way to try to dominate AI diffusion, even if the capabilities of your models lag behind the frontier. But in another sense, at some point, these models are getting really, really capable, including at doing things that the Chinese Communist Party might not like, like hacking their own critical infrastructure or getting around the Great Firewall.

And I just wonder, do you think that the authorities in Beijing are going to continue to promote open weight models? Or at a certain point, do you think that they will cap the release of open weight models at a certain capability just because of a loss of control concern they might have?

Chhabra: I think it's a really, really important question, Colin.

So first, let me just say: I very much agree with Sasha. I think a healthy ecosystem is definitely going to include open models, proprietary models, but I think there are at least three factors here, and one is what you just described, which is the need for control on the part of the CCP is really insatiable. And so it is hard to see that there does not come a point where they become concerned about the capabilities that are let out into the wild, including cyber capabilities, and I think we could think about biocapabilities coming soon as well.

I think second, as Sasha knows as well as I do, that the current position of the U.S. government is for the frontier, particularly for models that are less safeguarded, they need to be in trusted access programs where you really know the actor, you trust the actor given the potential for harm. And second, it's been that where models are general access but very capable, there need to be very, very strong safeguards that the government itself now is testing. So that's kind of the U.S. government position right now.

I think the final piece of this is we shouldn't think about this totally ahistorically. We have seen this movie before where China provides very, very significant subsidies to eat up market share, not working within a free and fair market, and then come in and in a predatory way go after all competitors.

And remember, for many of the technologies where we have seen that happen before, there hasn't been necessarily a dedicated Polit Bureau session to discuss what their global strategy should be. There has been with AI, and Xi Jinping has been very clear on how he thinks about AI and how important it is as a strategic technology as well. So we should assume that the same playbook we've seen over and over again in other strategic technologies is at work here as well.

Kahl: Both you and Sasha have mentioned these safety and security issues. So maybe let's dive deeper into some of the risks that people are thinking about. 

A lot was made earlier this year about the cybersecurity capabilities when Anthropic held back the release, initially, of the Mythos model and established this Project Glasswing to kind of go shields up before the model went into the wild.

Obviously, Sasha, OpenAI's GPT 5.6 is an extraordinarily capable model at a lot of things, including coding. Some people have said, basically, that your companies have essentially created a skeleton key to the internet, that we now have models that are so good at identifying vulnerabilities and exploits that they can hack into any web browser, any legacy software.

Sasha, I read a blog post from OpenAI that said you were testing some combination of GPT-5.6 Sol and a new pre-release model in what was thought to be a closed sandbox on its cyber capabilities, and it hacked its way out of the sandbox, escaped into the open internet, got its way into a Hugging Face server and tried to steal secret information that would allow it to cheat on the evaluations you were you were doing.

So look, these models appear to be very good at hacking and are increasingly slippery. 

What is what keeps you up at night? Is it these cybersecurity risks? Tarun mentioned biosecurity risks. I've heard people talk about the possibility of maybe a Mythos moment for bio in 2026. Is it the prospect of recursive self-improvement of models that become able to improve themselves and perhaps become increasingly autonomous and out of control?

What are the AI risks that you're going around the world, Sasha, talking to leaders and enterprises that you're most concerned about?

Baker: To a certain extent, it's a little bit of all of the above. Maybe just to talk about the Hugging Face incident first. It’s an example of reward hacking, essentially, by the model. The model was given a task and it was very determined to complete that task. And because of the information that it had in its possession, it knew that the repository with the answer key existed in this Hugging Face repository.

So, it's an example of model determination, I guess, if nothing else. And of course, some very significant capability. We're doing as you would expect and taking a hard look at our security parameters around some of these models and the containers that we keep them in to make sure that we're up-leveling that as the models become more capable.

And that's maybe the one item I would put on your list that you didn't already mention, which is I think we need a new paradigm about thinking about model safety and security for agentic models that can take action on your behalf because that changes the dynamics and there are lots of ways that that could go sideways, including inadvertently. You don't have to have a nefarious intent. If a model doesn't fully understand what its boundaries and its guardrails are, it can do something that you might not expect it to do in the course of completing a task that you did expect it to to complete. And I think a little bit of that is what we're seeing here.

So that's an area where I think we, you know, we collectively as an industry need to pay more attention and a little bit more research.

Cyber and bio are challenges because they're inherently dual use, right? There are a lot of things that we would want these models to enable. We want them to be able to help create cures for diseases that currently have no solutions. We want them to help vetted cyber defenders protect their perimeters. But we have to then have a way as responsible actors In the ecosystem of trying to prevent those same capabilities from landing in the hands of somebody who might use them to do something that we as a human species, as a population, wouldn't want to see them do, right? 

So that's the reason that I think both Anthropic — and not to speak for Anthropic, Tarun—but both Anthropic and OpenAI have invested so heavily in these trusted access type programs, whether it's our trusted access or Glasswing, and in coordinating that to make sure that we really know who is using these tools and for what purpose.

Kahl: Tarun, what's your assessment of the risks? And maybe just to connect it back to our previous conversation on U.S. and and China, do you assess that the risks are different between the closed models that you're releasing and the open models that China tends to be releasing?

Chhabra: I agree with Sasha. We spend a lot of time talking about all of the above and I think the alignment issues come into sharp focus when we have incidents like what Sasha just described and credit to OpenAI for sharing that with everybody in a timely way. We've tried to do the same thing when we've seen examples of similar deceptive behavior. I think the alignment challenges are really, really important and hopefully we'll all be talking about them more collectively.

I think on the China side, this goes back to your question, Colin, about whether we're gonna hit a certain threshold for them where they are more worried about capabilities being released into the wild.

For a while you could speculate that they felt like they were more protected behind the firewall and that we were more vulnerable from a cyber perspective, and that they had demonstrated that by supporting Vol Typhoon, Self Typhoon, Name Your Typhoon, implanting into our and allied critical infrastructure. But it may well be that they hit a certain threshold where they worry about their own security, too.

I think the pure technical challenge with safeguards is, as we know, when you have access to all the weights, they can be more trivially broken. And that's something that, obviously, the U.S. government itself has been concerned about when they've asked us to kind of share with them the results of our own testing on our proprietary models, and then wanted to, I think rightly and understandably and commendably, test them themselves, too.

Kahl: Let's pause on the alignment question because both you and Sasha have raised it. How should we think about alignment? When alignment was first coming into the discourse around AI, frankly, I think a lot of people had in their minds like the science fiction image of a rogue superintelligence that basically tries to kill or enslave us all, right? HAL 9000, Skynet, the Matrix.

But I think what we could also imagine is just really, really powerful AI agents that have a lot of autonomy to complete tasks that they were given that generate outcomes that are not aligned with human interests or values. Not because the model is evil, but just because there's something about the model that we don't understand, or it's reward hacking, or it's doing something that creates a non-aligned outcome, even if the model is not doing it like intentionally to be some Bond supervillain.

Tarun, how should we think about the alignment issue?

Chhabra: Our approach to this has been first, we should be investing heavily in it, and we've been doing it from the earliest days of the company. And we have leading researchers like Chris Olah on the case, and we've been building out that team in a very, very significant way.

I think what we have tried to do is to document the earliest cases, even when they seem minor, even when they seem potentially a bit more trivial, just to document that this is emergent behavior that we ought to be worried about because to your point earlier and to Sasha's point earlier, as we see agentic activity really proliferate and as agents take on more and more consequential tasks, the ways in which you could have misalignment could really compound in terms of the consequences.

That's something that I hope we can continue to work with not only our enterprise customers, but also we've heard lots of great questions and important questions from our government colleagues about, too. They understand the ways in which they're likely to expand and use agents and have the same questions: how do they ensure that their agents are behaving in the ways that they would expect of the most professional intelligence or defense officials where the work is currently being done by humans?

Kahl: You mentioned your interactions with government officials. Let me ask a question about where you think the Trump administration is headed on this.

Early on in the second Trump administration, they were not too keen on AI safety. Although the Trump AI action plan in the summer of 2025 did have a section on what they called AI security, which noted risks around cyber and bio and some other concerns.

They appear to have become much more concerned about AI safety and security in recent months. We've obviously seen them take some actions to hold back the deployment of Anthropic’s Mythos and Fable models. They've also, I think, asked OpenAI to limit the initial deployment of GPT 5.6. A friend and colleague of ours, Dean Ball, has suggested that the Trump administration is trending towards a de facto licensing regime, essentially, on Frontier AI. 

Where do you think they're headed? Where do you think the administration is headed in terms of its requirements to do some testing and evaluation and kind of kick the tires on these things before it lets you release them into the wild. 

Sasha, maybe start with you.

Baker: I mean there's definitely an evolution happening here, and we're seeing more government officials across a broad range of agencies taking an interest in sort of understanding that the most capable AI models really do have security and safety significance.

I will say, there is a consistency, a through line here though. I have been in this role here at OpenAI now for about two years. So, I started in the last administration; I continued in this administration. And I actually am having a lot of the same conversations, right?

Because when you talk about national security risk, which is really a lot of what we mean when we say safety. It's not the only thing we mean, but a big chunk of it is cyber, bio, CBRN, things that are kind of in the national security domain, we find that governments have been paying attention to that for a while. It's certainly risen in prominence, and I think is certainly more public now than it was before. But the gist of the conversations hasn't changed all that much.

So where are they going? I'm not sure. If you know, we would love to know. But what I can say is that our feeling is like it's inherently a good thing, and we welcome the government being involved in this space.

Both Anthropic and OpenAI have had long-standing voluntary partnerships with what's now called the KC and with the UK AC, the AI Safety Institute in the UK as well, in part because we do think that governments should have a role in understanding and evaluating what these models are and what they're capable of. And we want that to be an ongoing conversation. And as the models get better, we think that that conversation probably needs to continue to become more robust as well.

So whether Congress passes a law, whether the administration takes action on its own, whether this remains voluntary, I can't predict. I can tell you that we will continue to volunteer because we think it's the right thing to do.

Kahl: Whatever one thinks of the administration's policy shift, it does strike me that you all would benefit from some degree of transparency over the standards against which your models are being judged and also some process that's predictable. So that when you go to them, you kind of can plan around, okay, it's gonna be 30 days and we have to release the model to KC and give this version to NSA, and they're going to hold it to these standards and we'll send engineers to help, blah, blah, blah, blah.

Do you have a sense that they are moving towards a more predictable process instead of standards that you all can plan around?

Chhabra: I think so. I think we can kind of already see what the emergent regime looks like based on what is being asked of us right now.

They want pre-deployment testing. They want to be able to test the safeguards when a model is generally available. They want a say in what a trusted access program looks like. We probably also all want some protocols on what happens when there's an alleged jailbreak incident, because we can imagine scenarios in which people could try to exploit fears about that, including adversaries. And so we should all have a playbook for what that looks like.

So in each of those areas, it's in everyone's interest—including the government's interest—to have a predictable and transparent regime for what this looks like so everyone can prepare because on their side, at a minimum, they want to make sure they have the right capabilities, the right people in place, the right protocols in place to kind of handle all the incoming because we we move at a pretty fast pace in putting out new models, in sharing new capabilities, and sharing what new risks look like. And so we just have to partner together on that.

And we’ve been asked for a lot of input on what this should look like. And so we're working together with them on it.

Baker: Maybe just to foot stamp one thing Tarun said, because I think it's really important, which is about capacity. There is a need for more AI expertise inside the government, across the board, but particularly when it comes to doing these kinds of technical evaluations of frontier models. And I give a lot of credit to the administration for trying some really innovative ways of bringing some of that talent into government.

But that is an area where I think we as industry can do more to lean in and support those efforts because in order for this to work well, there needs to be common understanding on both sides. And in order to have that, you really do need the technical understanding that's resident in the KC. It's resident in a couple other places in the government right now. But wouldn't it be great if we could just like 10x that?

Chhabra: If I could just add to Sasha's point here.

I think there's some really, really good news here, which is sometimes there's a misconception that in order to bring the most talented folks in government who have expertise in model development or safety or alignment, you have to kind of pay them outsized sums that are the same as what they are earning in the private sector. And it's just not true. Our colleagues at OpenAI, at Anthropic, at Google DeepMind who are developing these models are deeply mission oriented.

And if given the opportunity to work in a space where they know they will have impact, they have folks who will listen to them, you will have plenty of folks volunteering to do this work, especially now that there's a pretty strong direction to take these risks seriously.

And I found that to be true in government as well. It's not a coincidence that we were able to actually impose the initial export controls on China a month before ChatGPT was actually released, anticipating kind of where things were headed. We had the benefit of really terrific experts who wanted to serve in government because they knew they could have that kind of impact.

I think if we kind of create the right opportunities for them, we can ensure the right impact, we can really bring the talent that we need into the government.

Kahl: Well, I think all of us believe that public service is super important and that brilliant people should be motivated to serve their country to keep it safe and prosperous and free, even if they don't make the salaries they're making in the private industry.

I do wonder . . . we've mentioned Google a couple of times . . . I think Google has put forward a policy suggestion of creating basically an external auditing entity that maybe would be funded by industry, but not obviously governed by industry. It might actually be able to recruit and pay people a little bit more that would essentially work alongside government to audit your models based on your own safety criteria.

Is that something Anthropic has also talked about, and then Sasha, is this something OpenAI has talked about, or do you think their proper place for this auditing to happen is in the government?

Chhabra: Our approach, Colin, has been to basically offer what we think are a number of viable models and what you just described, we think, is one of them. There are a number of avenues that you could pursue.

I think though, whatever path you pursue with some sort of external testing capacity, the government is always going to want to have the ability internally to test when they want to and need to, and I think that's a good idea. That may be when they feel like they actually need to verify something an outside entity has tested and provided, or it may be that they have their own tests, you know, which they don't necessarily want to share with an external body, and there may be national security reasons for that as well.

Kahl: And Sasha, does Open AI have a view on whether there should be an external auditing entity in addition to the government?

Baker: I think we're interested in the idea that Google has put forward. There are obviously some mechanics of it that would need to be worked out and the details would need to be figured out. But there are other examples of how similar paradigms work in other industries, right?

You could think about like FINRA and the FCC as one model of something like that where there's sort of a government oversight body and a government accreditation body, but then there is an industry monitoring mechanism that is independent of government.

And so we're interested in this. We're talking with Google. I think Anthropic is as well, and we’ll see where those conversations go.

The other thing that we're really interested in — and I know, Tarun, we’ve talked about this in the past — is building out more of that independent evaluation ecosystem because right now we all work with a number of the same independent evaluators who have the expertise and the data sets and the benchmarks that we use to evaluate our models.

But the truth is as the models get better, we need new benchmarks because those benchmarks are getting saturated. And those are time intensive and they are data intensive and they are expertise intensive to create. And so the more that we can collectively do to up-level that outside ecosystem, whether it's in collaboration with government, whether it's industry funded—we're all members of the Frontier Model Forum, which is the sort of safety-oriented frontier model industry association—there are lots of ways that you can kind of get at this.

But I do think that there's starting to be a prevailing view that we need certainty in this process. We need to be able to scale the process as the models scale, and that we need to be in constant coordination both with each other and with the government.

Kahl: Sasha, I want to tap into your Pentagon experience for a minute.

Obviously we've talked a lot about the AI risk side of the equation in the security space, but there are a lot of national security applications for AI with a lot of upside for national security. 

We've seen in the wars in Ukraine and the Middle East AI being used. It's fusing intelligence. It's helping enable battlefield management. There are increasingly autonomous drones being used, especially in Ukraine.

I wonder again, with your former Pentagon hat on, as you look at the landscape, where do you think the most promising national security applications for frontier AI models are right now?

Baker: Thank you for your question. Because first it gives me an opportunity to pitch something that we just put out, which is a National Securities Principles document. Colin, I know you've seen this and we know some of the folks who worked on it behind the scenes.

Kahl: Yeah, it's a good document.

Baker: It was a really intensive and I think thoughtful effort across the company to try to articulate in a clear and enduring way how we approach questions of using AI models in this space. And it's a document we're pretty proud of.

So you can find it on the internet. If your listeners want to read it, you can go to our website and I encourage that.

But to answer your question. I get really jazzed about this question because I am still sort of a Pentagon nerd at heart. I would say maybe two things.

The first is there's so much low-hanging fruit that is not stuff that people are thinking about or talking about every day, and it's frankly not that controversial.

The U.S. military is the biggest bureaucracy in the world. You're talking about three million people, HR, healthcare, logistics, audit. All of these things are incredibly data intensive and places where AI tools — and frankly things that we've done in other industries — could be applied to save money, to create, to improve people's lives, to make workflows more efficient. That is the table stakes, and we should have been doing that stuff yesterday.

And then beyond that, I think the area where AI tools for me show the most promise has to do with what they're best at, right? Which is helping people process enormous amounts of data and information.

And when you think about what a modern battlefield looks like, it is essentially a data-saturated environment. And so the more that models can do to help humans . . . because you know, we talked about human in the loop and wanting to retain human judgment over high consequence decisions, including the use of force. But the ways in which models I think can be used appropriately and responsibly to help humans make better decisions faster is an area where I think we've only begun to kind of scratch the surface. And so there's a lot I think that we could do there.

When I talk about this internally and I talk about this even with governments around the world, I think it was Colin Powell who said this, right? That you never want to send your military into a fair fight. You always want to equip them with the tools that are going to allow them to have the greatest chance of coming home safely. And that is, for me, principle number one and the reason why I'm here and the reason why I feel so passionately about making sure that we have these partnerships with government in the national security space.

Kahl: Yeah. When you and I were at the Pentagon, Kath Hicks, who was the deputy secretary, used to talk about AI as a means for decision advantage, which I think is very much along the lines of what you just talked about.

Tarun, reports suggest that Claude is part of the Maven Smart System AI platform that is being used by the U.S. military in its current conflicts.

What are the biggest national security applications as you see it?

Chhabra: I think as Sasha said, there's a ton to be done at the enterprise level. And sometimes the best way to do that is just for senior military leaders to hear from leaders in enterprise about how they're using things for all of the things that Sasha described.

But I think it's a very straightforward proposition. It's see the battlefield more clearly with more precision and more breadth than the adversary. There's just incredible power in doing that.

And having your adversary know that we can do that obviously has powerful deterrence value as well because it enables far more decision support and it enables much, much much speedier action as well.

I think one of the areas where we're looking now, and I know you know OpenAI is doing the same, is where can we also try to help make up for the deficit in manufacturing, particularly for the defense innovation base. And could we use frontier models now to catch up and maybe even leapfrog Chinese capabilities if we stay at the frontier?

Already the models show a lot of promise without much fine-tuning in supporting robotics operations, for example. And we think there's a lot more that can be done here in the defense manufacturing base. So we're really excited about that work and hope all the labs can contribute to that.

Kahl: Awesome. Well, look, you know, one of my favorite podcasts is Ezra Klein's podcast. And at the end of his podcast, he always asks the guest for three books. Which I always feel intimidated by, because even as an academic, I don't have the time to read a single book most of the time. So I'm in the habit of asking our guests for a single article.

So, what is one article from each of you—Tarun we'll start with you and conclude with Sasha—one article you might recommend that our listeners check out to either understand AI or some other aspect of the world in 2026.

Tarun, any suggestions?

Chhabra: I think I have to cheat with two. I think Daniel Kokotajlo's AI 2027 work and that of his colleagues has actually aged pretty well. Maybe they actually underestimated the pace at which AI would progress. But I think kind of capturing how a government would think about some of the risks is really important to revisit today. When it came out in draft in 2024, it was a little bit far-fetched for a lot of people, but if you read it today, it doesn't look that way anymore.

The other one is, actually, since I'm talking to a Stanford professor: one of my favorite classes I took was with Gavin Wright in the history department who taught American economic history. Maybe he's still teaching a version of that course. And one of the things we read was an article by Francis Thompson, which was “Nineteenth-Century Horse Sense.” So it’s like the rise and fall of horses in the Victorian British economy.

And one of the things he documents there is with the advent of the steam engine, you actually had increased use of horses at the endpoints because you just had these isolated channels otherwise and without using more horses—so it's a version of Jevon's paradox—you actually couldn't make much use of it.

But then you hit a cliff at some point when the horses no longer became economically as efficient. But there were so many social and political choices that had to be made along the way, and so I think it's useful to think about that analogy today.

Kahl: Sasha, any farm animals on your list?

Baker: I can't say I've read the article about horses. Although it sounds interesting!

I'm going to cheat in a different direction and I'm gonna recommend a documentary.

There's a documentary called Alpha Go. It's about the deep mind model that was able to win the Go competition.

And what I think is enduring about that moment is it's really about technological surprise and how humans adapt and react to growing machine capability. And so in that sense, there's like an interesting throughline to the moment that we're in now. I'm pretty sure it's still available on Netflix. So if you're like me and you spend all day staring at words on a screen or paper and you want to see moving images instead, that's where I would start.

Kahl: My recollection—tell me if this is wrong. But Go is one of the world's oldest games. It's really, really difficult to master. It was assumed that AI could never do it. And then, Alpha Go basically, I think it was Move 37, infamously, came up with a move that so stunned the world's best Go player that he quit. And so it's like, AI doing the impossible, perhaps.

Baker: That's a good note to end on, Colin. AI doing the impossible!

Kahl: Well, thank you so much, Sasha. Thank you, Tarun, for taking time out of your busy schedules to make us all smarter about AI and national security. Good luck with everything, and we hope to have you back on the pod at some point in the future.

You've all been listening to World Class from the Freeman Spogli Institute for International Studies at Stanford University. If you like what you're hearing, please leave us a review and be sure to subscribe on Apple, Spotify, or wherever you get your podcasts to stay up to date on what's happening in the world, and why.

Read More

Colin Kahl, Director of the Freeman Spogli Institute for International Studies, on stage with panelists at the May 5 event, "World Changing Technology in 2026"
News

FSI Scholars Examine AI, Biotech Advances, and Geopolitical Competition

At a May panel discussion, experts from across the institute assessed biotechnology's resurgence, the mental health effects of social media, and growing concerns about AI-enabled bioweapons.
FSI Scholars Examine AI, Biotech Advances, and Geopolitical Competition
Eyck Freymann on the World Class podcast
Commentary

Uniting America's National Powers to Prevent a War Over Taiwan

Eyck Freymann joins Colin Kahl on the World Class podcast to explain his plan to bring America's military strength, economic leverage, technological leadership, and diplomatic influence together into a single, coherent plan to curtail China's ambitions toward Taiwan.
Uniting America's National Powers to Prevent a War Over Taiwan
A panel of men sit at a long table on a stage.
News

China's Innovative Capacity Is Underestimated — and the Stakes Are Growing

SCCEI brought together leading China scholars this spring for its third annual China Conference under the theme “Understanding ‘DeepSeek Moments’ and China’s Innovation Ecosystem.” Conversation centered around the idea that the world’s prevailing frameworks for assessing China’s innovative capacity often underestimate it, and the consequences of that blind spot are growing.
China's Innovative Capacity Is Underestimated — and the Stakes Are Growing
Hero Image
All News button
1
Subtitle

The heads of national security policy at OpenAI and Anthropic join Colin Kahl on the World Class podcast to discuss how AI is changing national security strategies and the nature of U.S.-China competition.

Date Label
Display Hero Image Wide (1320px)
No
Paragraphs

Washington’s alliances are under immense strain. Many allies and partners are subject to increased threats from great-power adversaries, and they are coming to doubt whether they can rely on the United States. The response to these pressures is to rearm. Like the United States itself, U.S. partners across Asia, Europe, and elsewhere are building up their defense industrial and technological bases to improve their ability to project power, deter enemies, and prevail in a protracted conflict.

Continue reading at foreignaffairs.com

All Publications button
0
Publication Type
Commentary
Publication Date
Subtitle

America and Its Allies Must Pool Their Efforts

Journal Publisher
Foreign Affairs
Paragraphs

Escalating threats to undersea cable networks, stemming from gray-zone sabotage at vulnerable chokepoints, are receiving long-overdue attention from policymakers and the public. Recent incidents highlight the strategic vulnerability of this infrastructure, due to a lack of redundancy, limited repair capacity, and gaps in international maritime law. Despite attempts at multilateral cooperation through the G7 and the Quad, concrete actions are lagging. The US administration has not directly addressed this issue. To strengthen resilience, democracies must collaborate and invest in hardened cable designs, real-time monitoring and data sharing, routing diversity, regional repair hubs, and enhanced legal frameworks. They must work together to secure their lifeline for economic and national security and future digital-technology advancement.

All Publications button
1
Publication Type
Journal Articles
Publication Date
Journal Publisher
Texas National Security Review
Authors
Charles Mok
Number
Iss 3
Authors
Noa Ronkin
News Type
News
Date
Paragraphs

Across the world, populations are aging rapidly as people live longer and fertility rates continue to decline. Asia is at the vanguard of this demographic shift. The number of older adults (aged 60 and above) in the region is projected to triple between 2010 and 2050, reaching nearly 1.3 billion people. As Asian economies face this “silver wave,” helping older adults live safely and independently at home – a concept known as aging in place – has become a policy imperative.

At a recent webinar held during Stanford Health AI Week, the Asia Health Policy Program (AHPP) at Shorenstein APARC brought together experts from China, Singapore, and South Korea to share insights into the potential of health AI to allow older adults to enjoy healthy aging and avoid or postpone institutionalization. 

Moderated by Stanford health economist Karen Eggleston, the director of AHPP, the webinar featured Hongsoo Kim, a professor of health policy and aging at Seoul National University’s Graduate School of Public Health and director of its Artificial Intelligence Institute’s Center for AI in Health and Care; Xiaochen Ma, an assistant professor of health economics at Peking University’s China Center for Health Development Studies; and Tien Yin Wong, a physician-scientist-innovator and the senior vice-chancellor of Tsinghua Medicine and vice-provost of Tsinghua University, who has also worked and held senior leadership roles in Singapore and Australia as a practicing retinal specialist with a research portfolio on retinal diseases, ocular imaging, AI, and digital technology.

Get APARC event invitations and guest speakers' insights delivered to your inbox >



Here are six lessons from the front lines of Asia’s efforts to integrate AI into elderly health care and advance aging in place:

1. Adopt a Whole Systems Approach


In South Korea, the world's fastest-ageing society, automated systems like "CLOVA CareCall" – an AI-powered well-being dialer – conduct natural-sounding check-ins with solo-dwelling seniors, boasting a 96% response rate. Yet, Professor Kim emphasizes that checking in with people in need of health care is only half the battle.

If an AI flags an isolated senior at risk of depression, cognitive decline, or a physical abnormality, but the local community lacks the social workers or clinical pathways to intervene, then the health care system has failed.

“The question is not only whether AI can detect something, but how a health and care system acts on it,” she says. “Detection by itself changes nothing. A warning that no one follows up on helps no one. So the gap I care about is not the model’s cleverness itself. It is whether the system delivers.”

2. Solve the Entire "Care Cascade"


In rural China, traditional diabetic screening rates hover below 33%, leaving millions at risk of Diabetic Retinopathy (DR), a leading cause of blindness. Professor Ma shared how deploying an AI screening model successfully pushed screening rates past 85%.

The research team, however, discovered a glaring bottleneck: only 21% of high-risk patients actually followed up to receive sight-saving treatments. To fill this gap, Ma’s team designed an "AI Plus” model (v2.0) that integrates immediate, local-language counseling at the point of screening. To keep seniors healthy at home, AI solutions must address the entire clinical journey, from initial scan to final treatment.

“Many of the AI tools have been focused on diagnosis accuracy or validation rather than going downstream to the entire cascade of whether improved screening will translate into improved referral and the ultimate health outcomes,” says Ma.

3. Align with Local Workflows and Incentives


AI and other technology solutions for health often fail because they expect overworked care workers to adopt entirely new habits. Professor Ma noted that digital health interventions in rural China succeeded only when they integrated seamlessly into existing daily routines.

Instead of forcing clinicians to use complex new software, successful pilots utilized WeChat, the ubiquitous messaging app already open on every phone. Furthermore, the technology must align with the financial and professional incentives of frontline health workers. If an AI tool increases their administrative burden without simplifying their day or boosting their clinical efficiency, then it will remain unused.

4. Design Human-Centered AI for Health Equity


Professor Wong highlighted the ethical risk that AI tools will worsen, rather than reduce, health care disparities. This challenge is driven by the dynamics of “Inverse Care Law,” where AI disproportionately benefits the already advantaged, and the “Recursive Care Law,” where this inequality becomes a self-reinforcing cycle embedded in the system.

Because younger, more tech-savvy individuals generate more health data, AI models become better at serving them than the intended users of aging-in-place technologies. This creates a vicious cycle where the very tools designed to support aging populations end up marginalizing them. Governments must devise policies to mandate fair data coverage and usability, ensuring that AI serves society's most vulnerable members equitably, Wong stated.

Professor Kim noted that her team found that only about 38% of community care agencies in Korea have adopted AI and that the adoption rate varied sharply by region. In fact, districts with the greatest need may have the least access to these powerful tools. This challenge is a fundamental design gap rather than a technology gap, Professor Kim argues. To be genuinely equitable, a system must be built from the start to actively track who is missing and automatically route support back to them. This requires two  human-centered design key principles:

I. Universal by Default: The hardest-to-reach should not have to be the most persistent in navigating the technology.

II. Connected Across Sectors: Long-term care, social care, and health care must act as one integrated system rather than disconnected silos, each of which sees only part of the person’s needs.

5. Augment, Do Not Replace, the Human Touch


The panelists rejected the trope of robots replacing human caregivers. Instead, they view AI as an essential force multiplier for an overstretched workforce.

Whether it is South Korea’s deployment of 12,000 AI companion robots to combat senior isolation, or automated triage tools in clinics, the goal should be to offload administrative and routine tasks. This frees up human social workers and clinicians to do what they do best: deliver hands-on, empathetic care.

6. Value Real-World Outcomes Over Technical Novelty


Healthcare systems should prioritize rigorous, real-world case studies that prove actual clinical value, such as reduced mortality, lower rates of blindness, or fewer nursing home admissions, rather than celebrating high validation benchmarks in a laboratory.

To build robust future health AI systems, the experts concluded, the academic and tech sectors must also courageously publish and analyze their failed trials to understand what truly works in the chaotic reality of home-based care.

While AI holds immense promise for helping people grow old at home, “age tech” alone cannot solve the elder care crisis, the panelists agreed.


 

In the Media


Dr. Eggleston joined Jeffrey Snyder, host of the Broadcast Retirement Network, to discuss highlights and takeaways from the webinar discussion. Watch below:

Read More

Four elderly Chinese people sitting outdoors.
News

Asia's Aging Populations Drive Surging Disease Burden, Although Individual Health Improves

Across five Asian health care systems, rapid population aging drives up disease burden, particularly for chronic conditions, even as medical advancements improve outcomes for individual patients, according to a study co-authored by Stanford health economist Karen Eggleston.
Asia's Aging Populations Drive Surging Disease Burden, Although Individual Health Improves
A teenager is given blood test during a physical examination in Seoul, South Korea.
News

Income-Based Health Inequalities Persist in the US and South Korea, Though Universal Coverage Helps Reduce Disparities

South Korea achieves comparable clinical outcomes at lower per-capita spending than the United States, according to a new study. The co-authors, including Stanford health economist Karen Eggleston, find systemic income-based inequalities in health care access and utilization in both countries, albeit they are less pronounced under South Korea's universal health care system.
Income-Based Health Inequalities Persist in the US and South Korea, Though Universal Coverage Helps Reduce Disparities
Hero Image
Elderly woman's hand touching a robot's hand. Photo from a summit on AI's potential for empowering humanity.
All News button
1
Subtitle

Top aging and healthy policy experts from China, Singapore, and South Korea agree that helping older adults age at home requires addressing systemic health care bottlenecks rather than racing to build smarter AI models.

Date Label
Display Hero Image Wide (1320px)
Yes
Authors
News Type
Commentary
Date
Paragraphs

In its June 20, 2026, edition, the Nikkei Shimbun's Market Beat financial analysis column, titled US Stocks Attract Japanese Money, examines how the U.S. stock market is becoming a "major league" that is increasingly attracting capital from Japanese individual investors. Several factors drive this phenomenon:

  • Massive IPOs: Unprecedentedly large IPOs, as seen by SpaceX, are capturing global attention and investment.
  • Structural Advantages: U.S. markets are moving fast to include new, large companies in major stock indices, which forces index-tracking funds to buy shares and creates automatic demand.
  • 24-Hour Trading: U.S. exchanges like the NYSE and Nasdaq are moving toward near 24-hour trading to specifically capture Asian daytime investors, making it easier for them to participate.


The column describes how Japanese investors are increasingly bypassing domestic options to capture the immense growth driven by American AI, data infrastructure, and advanced technology IPOs. "Founders and early-stage investors prefer large-scale markets that offer price discovery, liquidity, and visibility. This is likely to accelerate the concentration of innovative companies in the United States," says Curtis Milhaupt, the William F. Baxter-Visa International Professor of Law and APARC faculty affiliate.

The article notes that the speculative nature of advanced tech sectors inevitably fuels market volatility and could pose risks for Japanese investors looking across the Pacific, urging Japanese investors to identify promising domestic companies.

Read More

A wave-shape graphics in Stanford cardinal red with text "Corporate National identity," a title of a working paper by the European Corporate Governance Institute.
Working Papers

Corporate National Identity

Corporate National Identity
Illustration of a semiconductor chip with the Japanese flag superimposed on top.
Working Papers

Japan’s Economic Security and the Semiconductor Industry

The Validity of the Revitalization Strategy
Japan’s Economic Security and the Semiconductor Industry
Hero Image
Portrait photo of Curtis Milhaupt and masthead of the Nikkei Shimbun.
All News button
1
Subtitle

Japanese capital is flowing rapidly into U.S. markets to back AI, tech IPOs, and data infrastructure.

Date Label
Display Hero Image Wide (1320px)
Yes
Paragraphs

A new front has opened in the U.S.-China competition in artificial intelligence: open-weight, local AI models. Until recently, the most capable AI models were too big and too costly to run anywhere but in giant data centers packed with expensive, specialized chips. But now these systems are rapidly migrating from the cloud to consumer hardware—including laptops and mobile devices—where they can answer questions, write code, and take actions on a user’s behalf without sending data to a remote server. Thanks to technological advances in both AI models and chips, the so-called open-weight AI models that increasingly underpin most local AI deployments are smarter and smaller than their predecessors and can be freely downloaded from the Internet, modified, and deployed without a centralized provider.

Continue reading at foreignaffairs.com

All Publications button
0
Publication Type
Commentary
Publication Date
Subtitle

How to Counter Beijing’s Unauthorized “Distillation”

Journal Publisher
Foreign Affairs
News Type
News
Date
Paragraphs

The Stanford Deliberative Democracy Lab, based at the University’s Center on Democracy, Development and the Rule of Law, today released the findings from two national Community Forums on the evolving expectations around privacy and governance of AI-powered wearable devices. In collaboration with Meta, the forum engaged a representative sample of 550 participants — 300 from the United States and 250 from India — to solicit people's perspectives on user controls and societal expectations. The Community Forums were conducted as national Deliberative Polls.

As AI wearables see rapid adoption worldwide, understanding public attitudes can help to ensure these technologies are developed and deployed responsibly. Three key themes emerged from the forum. Participants in both the U.S. and India indicated the highest levels of support for users having controls over when their wearables passively process environments and actively capture data. U.S. participants consistently favored individual agency over how wearables are used, both in public and private settings, whereas in India, there was a slight preference for governments to decide wearable usage rules in public spaces. Additionally, U.S. participants supported workplaces and schools having the primary authority to decide how AI-powered wearables should be used in those environments, while Indian participants also saw a significant role for governments in these settings.

The forum also revealed important nuances in public perspectives. For example, participants expressed a preference for AI wearables that are tailored to cultural and regional contexts, rather than standardized global designs. There was also broad support for AI agents capable of responding to emotional cues, underscoring the public's desire for personalized, human-centric wearable experiences.

"This global forum provided invaluable insights into how the public's expectations around privacy and governance of wearable AI are evolving," said Alice Siu, Associate Director of the Stanford Deliberative Democracy Lab. "The findings will be essential for policymakers, technology companies, and other stakeholders as they work to ensure these powerful technologies empower users while respecting fundamental rights."

Read More

Human finger touching a screen with AI agent
News

People Shaping AI: Groundbreaking Industry-Wide Forum Invites Public Input for the Future of AI Agents Through Public Deliberation

In an unprecedented collaboration, Stanford's Deliberative Democracy Lab has spearheaded the first-ever Industry-Wide Forum, a cross-industry effort putting everyday people at the center of decisions about AI agents.
People Shaping AI: Groundbreaking Industry-Wide Forum Invites Public Input for the Future of AI Agents Through Public Deliberation
Back view of crop anonymous female talking to a chatbot of computer while sitting at home
News

Meta and Stanford’s Deliberative Democracy Lab Release Results from Second Community Forum on Generative AI

Participants deliberated on ‘how should AI agents provide proactive, personalized experiences for users?’ and ‘how should AI agents and users interact?’
Meta and Stanford’s Deliberative Democracy Lab Release Results from Second Community Forum on Generative AI
Chatbot powered by AI. Transforming Industries and customer service. Yellow chatbot icon over smart phone in action. Modern 3D render
News

Navigating the Future of AI: Insights from the Second Meta Community Forum

A multinational Deliberative Poll unveils the global public's nuanced views on AI chatbots and their integration into society.
Navigating the Future of AI: Insights from the Second Meta Community Forum
Hero Image
Close-up of smart glasses showing augmented reality interface while held by a woman in a modern urban environment.
Getty Images
All News button
1
Subtitle

National community forums in the U.S. and India highlight differences in preferences for privacy, user control, and governance of emerging technologies.

Date Label
In Brief
  • Stanford’s Deliberative Democracy Lab convened national forums in the U.S. and India to examine public attitudes toward AI-powered wearable devices.
  • Participants in both countries strongly supported user control over data collection, with differences in preferences for government and institutional oversight.
  • Findings highlight demand for culturally tailored designs and personalized, human-centered AI features as adoption of wearables grows.
Display Hero Image Wide (1320px)
Yes
Subscribe to Technology