Microsoft

Agent Lightning v1.0: A 3,500-Line Lightweight Agentic RL Framework for Training Agents with Real Harnesses

Microsoft Research - Wed, 10/07/2026 - 18:00
At a glance
  • Harnessed Agentic RL: Microsoft Research Asia introduces a training paradigm in which the same agent harness used in deployment participates directly in reinforcement learning, removing the need to reimplement the agent inside the training framework.
  • Lightweight by design: Agent Lightning v1.0 delivers a complete agent RL control plane in roughly 3,500 lines of code.
  • Native Kubernetes support: agents run as standard Kubernetes jobs on self-managed clusters, cloud Kubernetes, or local infrastructure, with no dependency on paid commercial sandbox services.
  • Data-efficient training recipe: an end-to-end coding agent pipeline raised Qwen3.5-9B from 41.8% to 56.4% Pass@1 on SWE-bench Verified, a 14.6 percentage point gain, using only about 6,000 training samples based on open sourced dataset.

AI agents have evolved from single models to complex full-stack systems built from models, tools, and execution environments. Their capabilities increasingly depend on the agent harness that coordinates them from outside the model. Reinforcement learning (RL) is an approach where AI systems learn through trial and error, guided by rewards and penalties for their actions. RL can make those agents better, but most agent RL systems require developers to reimplement the agent inside the training framework. That is costly, and it means the agent being trained is not quite the agent that gets deployed.

To address this, researchers at Microsoft Research Asia have introduced the Harnessed Agentic RL training paradigm and open-sourced a fully rebuilt Agent Lightning v1.0 (opens in new tab). Compared with the original, Agent Lightning, v1.0 puts more emphasis on staying lightweight, on integrating with real harnesses, and on a complete, reproducible agent RL training pipeline.

Agent Lightning v1.0 was rebuilt around Harnessed Agentic RL, with key improvements:

  • Lightweight: the entire framework is about 3,500 lines of code. Agent Lightning v1.0 implements a complete Harnessed Agentic RL system in a codebase that is small and clear enough to understand, modify, and extend.
  • Training on a real agent harness: agents reach the model through the large language model (LLM) proxy in Agent Lightning v1.0, leaving existing harness code unchanged.
  • Native Kubernetes support: agents run directly as Kubernetes jobs, without external commercial sandbox services. Self-managed clusters and local infrastructure alike can support rollouts at scale.
  • A complete coding agent training example: an end-to-end pipeline built on Qwen3.5-9B raised Pass@1 on SWE-bench Verified from 41.8% to 56.4%, an absolute gain of 14.6 percentage points, using only about 6,000 training samples.
The limits of traditional agentic RL

Traditional agentic RL assumes the training framework owns the interaction loop with the environment. In a ReAct-style loop, the model generates an action, the environment returns an observation, the observation is appended to the context, and the model generates the next action, so the whole rollout maps onto one continuous token trajectory. Early RL systems such as verl, AReaL, and slime were built this way, which meant training an agent required rebuilding its loop inside the RL framework.

Real harnesses have outgrown that assumption. Coding agents such as mini-SWE-agent, OpenHands, OpenCode, Claude Code, and Codex each bring their own context management, tool protocols, execution logic, and dependencies, as do general-purpose agent systems. Rebuilding one for training is expensive, and the rebuilt agent may no longer behave in the same way as the deployed agent.

Agent Lightning takes a different route. It places an LLM proxy between the agent and the model. The agent continues to run as before: simply point the endpoint that previously called the model API at Agent Lightning, and the training framework can observe and record its model calls. In v1.0, the researchers go further and formally define this paradigm as Harnessed Agentic RL: whichever agent harness is used in deployment is the harness that takes part directly in reinforcement learning during training (Figure 1).

Figure 1. Traditional agentic RL compared with Harnessed Agentic RL. In traditional agentic RL, the training framework manages the environment and the agent loop. In Harnessed Agentic RL, the harness manages both. Four challenges in training with real harnesses

A core difference between Harnessed Agentic RL and traditional agentic RL is that the environment interaction loop is handled by the agent harness rather than the training framework. The training system can only observe a series of LLM request and response pairs, so a single rollout may be split into a variable number of training samples. This brings four key challenges:

  • Retokenization and sample merging: Harnesses keep context as text, but RL training needs the token IDs sampled during the rollout. Passing text through the chat template and tokenizer again can shift token boundaries, so adjacent calls cannot always be merged into one sample.
  • Advantage calculation: Retokenization, subagents and context summarization can split one rollout into several samples. Computing baselines and advantages directly at the sample level causes rollouts that produce more samples to be counted repeatedly, which alters the original statistical relationships at the rollout level.
  • Loss normalization: Averaging loss by sample count gives more weight to rollouts that produce more samples. Since sample count is often just a product of harness behavior, loss normalization must also avoid being distorted by it.
  • Training backend scheduling: Sample count and length are known only after the harness finishes, while GPU counts and data/tensor parallel configurations are usually fixed. The backend has to map a variable workload into fixed resources.

video series

On Second Thought

A video series with Sinead Bovell built around the questions everyone’s asking about AI. With expert voices from across Microsoft, we break down the tension and promise of this rapidly changing technology, exploring what’s evolving and what’s possible.

Explore the series Opens in a new tab Building a complete agent RL control plane in 3,500 lines of code

In system design, Agent Lightning v1.0 treats simplicity as its first principle. The entire framework is about 3,500 lines of code, with three core components: the API Gateway, the Rollout Controller, and the Customized Trainer (Figure 2).

The API gateway stores rollouts, models, and events, and serves as an OpenAI-compatible LLM proxy. It links every model call from the harness to its rollout and records the prompts, responses, and log probabilities that training needs. The rollout controller starts and manages agent execution, either as local processes or as standard Kubernetes jobs, keeping agent execution separate from the trainer. The customized trainer, built on verl, creates rollouts, waits for them to finish, collects samples, and assembles the final training samples through a sample adapter. As a result, for an existing agent harness, simply pointing the model endpoint at the Agent Lightning proxy is usually enough to connect quickly to RL training.

Figure 2. The Agent Lightning v1.0 system architecture showing the API gateway, rollout controller, and customized trainer. Collocated async RL

Rollout times vary widely across agents. Synchronous RL waits for the slowest agent in a batch and leaves GPUs idle, while fully asynchronous RL raises utilization but needs separate GPU pools for rollout and training. In response, Agent Lightning v1.0 introduces Collocated Async RL, which lets rollout and model updates share the same set of GPUs.

Once the system has collected enough rollouts, the update begins: the API Gateway pauses accepting new requests and waits for requests already in progress to finish, and rollout resumes after the update completes. The entire state transition is transparent to the external agent harness. In experiments, this approach achieved about a 2x end-to-end speedup over synchronous RL while using fewer GPUs than conventional asynchronous RL (Figure 3).

Figure 3. Synchronous RL, asynchronous RL, and Collocated Async RL compared. Collocated Async RL raises utilization while occupying fewer GPUs. Running agents on Kubernetes

Collecting enough rollouts means running many agents at once, which consumes substantial CPU, memory, and compute resources. Other Harnessed Agentic RL frameworks often host those agents on commercial sandbox services such as Modal Sandbox or E2B, where cost climbs quickly with scale. Instead, Agent Lightning v1.0 runs them as standard Kubernetes jobs, reusing existing self-managed clusters, cloud Kubernetes, or local infrastructure (Figure 4). Existing compute resources are used more efficiently, large rollouts cost less, and the whole pipeline stays open source and reproducible.

Figure 4. The Rollout Controller in Agent Lightning v1.0 provides native Kubernetes support, running agents directly as standard Kubernetes jobs. 6,000 training samples, a 14.6-point performance gain

To test the approach, researchers built a full pipeline on SWE-smith, mini-SWE-agent, and Qwen3.5-9B, covering data cleaning, environment construction, reward-hacking safeguards, and RL training. The training set holds about 6,000 samples and needs no large-scale compute. RL training alone raised Qwen3.5-9B from 41.8% to 56.4% on SWE-bench Verified, a gain of 14.6 percentage points.

The coding agent experiments further confirm the earlier analysis of two challenges: advantage calculation and loss normalization. Compared with sample-level handling, rollout-level advantage combined with rollout-level normalization achieves a higher validation reward and keeps policy entropy more stable during training (Figure 5).

Figure 5. Pass rate and policy entropy for Qwen3.5-9B on the SWE-smith validation set. Technical report GitHub project Opens in a new tab

The post Agent Lightning v1.0: A 3,500-Line Lightweight Agentic RL Framework for Training Agents with Real Harnesses appeared first on Microsoft Research.

Categories: Microsoft

Forecasting space weather risks on power grids

Microsoft Research - Wed, 09/30/2026 - 18:00
At a glance
  • End-to-end forecasting: A machine learning pipeline uses forecast-time solar-wind information to generate location-specific risk estimates for 66,935 substations in the continental United States. 
  • Physics and place: The system combines Auroral Electrojet (AE) and Disturbance Storm Time (Dst) forecasts with local latitude, geology, and ground conductivity. 
  • Advance warning: The pipeline detected nearly 80% of major space-weather events during the evaluation period and can warn grid operators 30 to 60 minutes before a specific risk appears. 
Forecasting a threat to critical infrastructure

During the May 2024 geomagnetic storm, utilities across North America prepared for possible impacts as auroras extended far beyond their usual range. The storm would degrade GPS accuracy and satellite operations, impacting farm operations in North America and Europe. Modern society depends on reliable electric power, yet extreme space-weather events can induce currents in transmission networks that damage equipment and increase operational risk. The challenge is not only knowing that a storm is approaching, but estimating when and where its effects could be most severe with enough warning for grid operators to respond.

During my summer internship at Microsoft Research, I developed a machine learning system that forecasts space-weather risk across 66,935 substations in the continental United States. The system measures how the sun’s activity can affect the earth’s magnetic field and, ultimately, the power grid.  It combines solar-wind observations, forecasts of the Auroral Electrojet (AE) and Disturbance Storm Time (Dst) indices, physics-informed constraints, local geological conductivity, and grid-infrastructure data to produce location-specific risk estimates 30-60 minutes ahead of potential impact. 

Building an end-to-end prediction pipeline

Space-weather prediction is difficult because it spans several coupled systems. The solar wind changes rapidly, its interaction with Earth’s magnetosphere is irregular, and the resulting ground effects depend on local conditions. Regions with resistive bedrock can experience stronger geomagnetically induced currents (GICs) than regions with more conductive geology. Transmission-line orientation, latitude, and other power-system characteristics further influence the exposure of individual assets.

The pipeline addresses this complexity in three stages, shown in Figure 1. First, solar-wind measurements from the L1 Lagrange point are used to generate forecasts of the AE and Dst indices, while geological conductivity and location features are assembled for each substation. Second, a gradient-boosting model combines these forecasted and location-specific inputs to estimate dB/dt, the rate of magnetic-field change associated with GIC risk. Finally, the resulting predictions are converted into location-specific risk estimates and aggregated into a continental risk assessment.  A system of 50 AI agents helped explore features, validation strategies, and model configurations across the pipeline. Only public data sources were used, included NASA OMNI and NASA-aggregated Kyoto World Data Center data, INTERMAGNET and U.S. Geological Survey magnetometer observations, and GridSFM-derived grid data. 

Figure 1. End-to-end forecasting pipeline from L1 solar-wind measurements to location-specific risk estimates across 66,935 U.S. substations. Evaluating geomagnetic forecasts and infrastructure risk

The AE predictor was designed to forecast rare, high-intensity geomagnetic activity that drives infrastructure risk. In the 2020-2026 evaluation period, the model produced forecasts spanning nearly the full observed range of AE activity and outperformed several empirical solar-wind-based approaches. The Dst predictor provided an additional signal describing large-scale geomagnetic storm strength. During the most geomagnetically active periods of the 2020-2026 evaluation period, the machine-learning model outperformed the Burton equation on 62.2% of individual hours. The model also produced a substantially wider prediction range than Burton-style approaches and improved severe-event detection in the end-to-end forecasting system by 1.2 percentage points when combined with AE forecasts.

PredictorResultBaseline comparisonImpactAE410.2 nT RMSELower RMSE than empirical and solar-wind-only baselinesCaptures extreme-event activity used by the GIC forecasting systemDst7.2 nT RMSEOutperformed Burton on 62.2% of peak-activity hours and provided wider prediction range+1.2pp severe storm detection when integrated into final GIC system

The GIC risk stage was evaluated differently. There is no equivalent widely deployed operational system that provides a direct industry benchmark for this calculation, so the machine learning model was compared with simple linear regression. As Figure 2 shows, the system achieved detection rates of 76.5% for major events (≥10 nT/min), 81.2% for severe events (≥20 nT/min), and 64.1% for extreme events (≥50 nT/min). False-alarm rates increased with storm severity, reflecting the trade-off between missed events and cautious alerts. Performance varied by latitude, with the highest detection rates at northern stations where geomagnetic activity is strongest.

Figure 2. Detection rates across storm-severity thresholds for regionally different substations, with a comparison between the average POD for each severity tier. Translating predictions into continental risk assessments

The final stage translates predicted geomagnetic activity into location-specific estimates of dB/dt, the rate of magnetic-field change associated with GIC exposure. Rather than issuing a single alert for the continental United States, the system combines storm conditions with each substation’s latitude and geological factor. This produces continuous risk estimates that can distinguish lower-risk locations from areas where resistive geology can amplify ground-level effects.

Figure 3 illustrates the output for a representative major-storm scenario. The map is a demonstration of the model’s continental-scale output, not a record of a live operational event. It shows how a grid operator or planner could move from a broad space-weather warning toward a more targeted view of which locations may warrant closer analysis. The current pipeline produced estimates for all 66,935 substations in approximately 333 milliseconds during measured inference, allowing many scenarios to be evaluated quickly.

Figure 3. Demonstration of a continental GIC risk assessment under a representative major-storm scenario. Colors indicate modeled risk levels across 66,935 substations. Implications and looking forward

This work demonstrates how physics-grounded machine learning could support more specific and timely assessment of space-weather exposure. Earlier, location-specific information could help utilities prioritize engineering review and consider targeted protective actions, such as adjusting reactive-power reserves or temporarily reconfiguring parts of the network. Further validation with utilities and operational data would be needed before the system could be used in grid operations.

The project also connects with broader Microsoft Research work on AI for power systems. The open grid-data pipeline provided realistic U.S. transmission models, while GridSFM applies deep learning to AC optimal power flow for fast scenario analysis. Together, these efforts point toward richer planning and resilience workflows that combine hazard forecasts, grid topology, and power-flow analysis.

  • Extend forecast horizons: Explore temporal-transformer approaches that capture longer-range patterns in solar-wind data and move beyond the current 30-60-minute window.
  • Scale internationally: Adapt the system to additional regions while accounting for different geological conditions and grid topologies.
  • Integrate with grid operations: Evaluate how forecasts could support existing decision-making and engineering-review workflows before considering higher levels of automation.
  • Provide transformer-level risk: Add asset-specific characteristics to move from substation-level estimates toward more granular assessments of critical equipment.
Related Microsoft Research work GridSFM: A new, small foundation model for the electric grid Building realistic electric transmission grid dataset at scale: a pipeline from open dataset Acknowledgements

Special thanks to my mentors Weiwei Yang and Spencer Fowers for their guidance throughout this project, and to Amber Hoak, Weishung Liu, and Andrea Britto for their editorial and technical feedback.

Opens in a new tab

The post Forecasting space weather risks on power grids appeared first on Microsoft Research.

Categories: Microsoft

Introducing Quine: An AI research system designed for the complexity of biology

Microsoft Research - Tue, 09/29/2026 - 16:00
At a glance
  • Quine (opens in new tab) is a research effort to create a multimodal world model of biology and an interactive harness connecting models, scientific tools, literature, and researchers. 
  • In collaboration with researchers at the Broad Institute of Harvard and MIT, we have used this system to prioritize compounds predicted to drive therapeutic tumor-state shifts and validated several top-ranked candidates across multiple wet-lab assays. 
  • The Quine Fellows program (opens in new tab) will give a cohort of scientists access to the system and an opportunity to accelerate their own research and provide scientific feedback.  
  • Quine is experimental research technology intended only for research, not clinical or medical use, and its outputs may be incomplete or inaccurate and require review by qualified researchers and appropriate scientific and experimental validation. As the technology matures, we expect to expand access through products like Microsoft Discovery (opens in new tab).

For more than two decades, Microsoft Research has worked at the intersection of computation and biology. Our research has spanned immunology, virology, genomics, biomedical imaging, cell biology, and protein engineering. That work has produced foundational methods, new science, and technology that reached the clinic, from rare and infectious disease diagnosis to cancer biomarker detection.

Across that work, one lesson has become increasingly clear: biology does not divide itself into the neat boundaries our models and tools often do. Genes influence proteins; proteins interact within cells; cells organize into tissues; and experiments continually reshape what scientists know and what they choose to ask next. Making progress on the hardest biological questions therefore requires more than increasingly capable models of individual datasets or tasks. It requires systems that can connect knowledge across scale and modalities, reason about experiments and evidence, and participate in the iterative process through which science advances. 

Today, Microsoft Research is introducing Quine (opens in new tab), a research effort designed to work across those boundaries, reflecting our long-term vision for a discovery system that evolves through scientific use. Quine brings together a world model of biology with a harness that connects scientific tools, literature, the wet lab, and the researchers using them.

The limits of experimentation

Even as experimental techniques have improved and wet-lab throughput has increased, biology remains fundamentally constrained by time and complexity. Nature cannot be rushed, nor can it be derived from first principles. Experiments are slow, iteration cycles are long, and many of the most important questions involve interactions, combinatorial design spaces, and downstream effects that are simply too large to explore experimentally alone.

At the same time, advances in large-scale machine learning, particularly the emergence of general-purpose foundation models and reasoning models that can iteratively work through problems, suggest a new possibility. These systems are beginning to demonstrate capabilities beyond pattern recognition: integrating information across domains, reasoning over abstractions, and supporting iterative problem-solving. Just as importantly, many of the techniques developed for human language have proven remarkably adaptable to aspects of biology, enabling models to learn representations of biological systems across diverse data types and scales.

This raised a provocative question for us: what would it take to build a world model for biology? By world model, we mean a system that can represent the state of a biological system, predict how that state will evolve in response to interventions, and reason over the consequences of those interventions multiple steps into the future. Such a system would not replace experimentation. Rather, it would allow us to use computation to explore, propose, rank, and prioritize potential paths forward before committing scarce laboratory resources. Our north star is a future in which scientists, models, and experiments operate in a continuously accelerating loop.  

What we’ve built

Quine is our first step toward that vision. It includes both a world model of biology and a harness that connects the model to orchestration and reasoning models, scientific tools, the literature, and the teams of scientists using them.  

Figure 1: Quine’s two frontier components — a world model of biology and an interactive harness — sit inside a loop that begins and ends with the scientist. A question becomes a set of proposals, proposals become designs worth testing, and the experiment returns measurements that sharpen both the scientist’s next question and the model itself. 

Core to this effort is the world model, which learns shared representations across biological modalities and scales, including sequence, structure, function, cellular state, and imaging data. By training across these modalities jointly, Quine can use evidence from one modality to inform predictions in another, capturing relationships that would otherwise remain siloed or lost when separate single-domain specialist models are orchestrated by the system. This multimodal design reflects the reality of biology: understanding any one level requires context from the others. Rather than diluting performance, we find that learning across these connected representations strengthens it, enabling the model to generalize more effectively and to support a broader range of biological reasoning tasks.

 Figure 2: Biology is multiscale and multimodal. (A) Biological evidence stacks from molecules to cells to tissues to patients, with scientific language running underneath all of it. (B) The modalities Project Quine trains on — genomics, proteins, chemistry, RNA and cell state, bioimaging — each observe a different span of that range, and they overlap. Reading down is the causal path from genotype to phenotype; reading across is physical scale. Because the model learns these representations jointly rather than orchestrating separate specialists, evidence at one level can inform predictions at another. 

By integrating this model in a unified system that facilitates computational experimentation and learns from iteration, researchers can generate predictions across modalities and reason about the consequences of biological interventions before they are tested in the lab. Critically, the world model doesn’t need to be perfect—indeed, it will never perfectly model biology—it simply needs to usefully inform experimental design. 

What this looks like in cancer biology

A concrete example comes from our work on pancreatic ductal adenocarcinoma (PDAC), the most common form of pancreatic cancer and one of the most challenging to treat. In collaboration with researchers at the Broad Institute of MIT and Harvard, we have spent years developing and applying patient-derived ex vivo models to investigate a longstanding hypothesis: that tumor behavior and drug response depend not only on genetics, but also on transcriptional cell state. In PDAC, tumor cells can occupy different cellular states associated with how they respond to treatment.  

With Quine, we have now put that hypothesis into practice, exploring whether non-genetic features of cancer, particularly cellular state, can serve as actionable therapeutic targets. We used Quine to predict and prioritize thousands of compounds based on their potential to shift tumor cells between therapeutically relevant states. In wet-lab studies focused on the transition from classical to basal cell states, Quine’s highest-ranked compounds produced the largest intended shifts across experimental assays. The entire process—from rapidly narrowing the compound search space to prioritizing a handful of promising candidates to be validated in the lab—took just one weekend, potentially saving months of experimental work and significant research costs. Notably, some of the strongest effects came from compounds with unexpected mechanisms of action, offering early evidence that AI can uncover new opportunities for drug repurposing and discovery. 

While the reverse state transition proved more difficult (from basal to classical cell states), Quine predicted that available compounds would have this weaker effect. More notably, the experiments revealed something we did not fully anticipate: Quine predicted that several compounds would consistently move cells toward a distinct third phenotype—an observation borne out in the lab—suggesting that the pancreatic cancer cell-state landscape is richer than a simple classical-basal axis. The experiments not only tested the model’s hypotheses but also generated new ones. 

Figure 3: Quine-guided exploration of pancreatic cancer cell states. In pancreatic ductaladenocarcinoma (PDAC) cell lines, top-ranked compounds produced the largest classical-to-basal transcriptional shifts as validated by wet lab experiments. Weaker reverse shifts and unexpected movement toward an additional phenotype were also consistent with Quine’s predictions.

Our continued work here will use newly integrated RNA datasets and tasks to represent this richer landscape, strengthen state-transition predictions, and provide calibrated confidence estimates that help scientists prioritize the most promising hypotheses for wet-lab testing. And this is exactly the feedback loop Quine was built to support. Each round of experimentation improves our understanding of biology, and each new insight can become training signal for the next generation of the system.  

Learning through discovery – an invitation

We created Quine to improve how we do science ourselves. As the system evolved, we embedded it in our ongoing scientific programs. Every research question, experimental result, and unexpected finding became an opportunity to refine the models, improve the workflows, and better understand where AI could meaningfully accelerate our research. Over time, the system became central to how we approached discovery across multiple domains, including in cancer biology, protein engineering, genomics, and bioimaging.

We’d like to continue building on that foundation by introducing Quine at this early stage. Quine grew out of our own research, but we don’t want its continued development to happen in isolation. Many of the questions that will determine where it goes next are not ones we can answer alone; they require the expertise, creativity, and perspectives of the broader scientific community. For this reason, we’re opening the Quine Fellows program to put the system directly into the hands of scientists working at the frontier of biology and medicine.

Applications for the first Quine Fellows cohort open now

Building responsibly will always come first. We believe progress in AI and biology must go hand in hand with safety, security, and responsible stewardship. We are taking a deliberate, phased approach to Quine’s development and access, with initial availability limited to the Quine Fellows program and select research collaborations. Ongoing internal review and built-in safeguards will help us maintain appropriate oversight as the system evolves. As the technology matures, we expect to expand access through products like Microsoft Discovery (opens in new tab).

It’s tempting to frame progress in AI for biology around benchmarks and leaderboards. Those matter, but the real test of Quine is what happens when it encounters science as it’s actually practiced. Can it be useful when evidence is incomplete? When facing questions genuinely never asked before? Can it help make the path from questions to discovery shorter?

Ultimately that’s the standard we care about. The ambition is for Quine to recede into the background of scientific practice, while the science moves faster in the foreground. The protagonists are not the model or the platform. They are the scientists, the experiments, and the discoveries that follow. 

Opens in a new tab

The post Introducing Quine: An AI research system designed for the complexity of biology appeared first on Microsoft Research.

Categories: Microsoft

One year in: How Microsoft Research Asia – Singapore is advancing research, partnership and talent for real-world impact

Microsoft Research - Mon, 09/28/2026 - 23:00

On July 24, 2025, Microsoft Research Asia – Singapore (MSRA – Singapore) opened its doors as Microsoft’s first research lab in Southeast Asia. The launch built on more than two decades of collaboration between Microsoft Research Asia and Singapore’s universities, research institutes, and government agencies. It marked a shift to a closer, more integrated relationship aimed at supporting Singapore’s innovation ecosystem and advancing the development and beneficial use of AI.

“When we established Microsoft Research Asia — Singapore one year ago, our vision was not simply to open another research lab. It was to build a bridge connecting Microsoft’s frontier AI research with Singapore’s world-class universities, forward-thinking government agencies, and dynamic industry ecosystem. One year in, what gives me the greatest confidence is not just the research we have produced, but the depth of partnership we have built. That is the foundation on which lasting impact is built.”

— Dr. Lidong Zhou, corporate vice president at Microsoft and managing director of Microsoft Research Asia

Established in 1998, MSRA has a presence in Beijing, Shanghai, Vancouver, and Tokyo, along with Singapore. The Singapore lab strengthens the collaboration with local partners and extends the scientific and societal impact of Microsoft’s global research community. 

Over the past year, the Singapore lab has grown its team, deepened partnerships, and explored how frontier AI research can create real-world value. Guided by its mission to advance frontier AI for humanity, its work spans four pillars: next-generation AI models and agentic systems, domain-specific AI for real-world impact, AI-native research practices, and ecosystem and talent development.

From frontier research to real-world impact – and back

Research has its greatest impact when it moves beyond the laboratory into practice: in health and medicine, enterprises, and the daily operations of organizations solving real problems. MSRA – Singapore aims to do just that through a two-way exchange: partners bring domain expertise, real-world data, and operational challenges, while the lab contributes frontier research and engineering capabilities. Together, they translate fundamental advances into practical solutions, while surfacing new questions that feed back into research in a kind of research-to-impact flywheel. Advances in AI foundations enable more capable models and agentic systems that can reason, plan, use tools, and collaborate with people. These capabilities create new opportunities for AI-native innovation. In turn, real-world applications reveal limitations and generate new questions, feeding insights back into foundational research and sustaining a continuous cycle of discovery, application, and learning.

Healthcare offers an early example of the lab’s research-to-impact model. Research on multimodal healthcare AI and self-evolving diagnostic agents is deployed and extended through partnerships across Singapore’s healthcare ecosystem. The lab is collaborating on multimodal and agentic AI approaches to support clinical decision-making, unlock insights from complex medical data, and advance research in disease diagnosis, risk assessment, and personalized care, with the longer-term goal of enabling meaningful real-world use.

Beyond healthcare, the lab is working with organizations across financial services, education, and technology to explore how frontier AI can address sector-specific needs and operational challenges.

“Over the past year, the lab has become an integral part of both Microsoft’s local ecosystem and Singapore’s broader innovation community. Its progress demonstrates the value of establishing a sustained research presence in Singapore.”

— Wee Luen Chia, managing director, Microsoft Singapore

Singapore has made AI a national priority through its National AI Strategy 2.0 (opens in new tab), which focuses on deploying industry-transforming AI, advancing frontier research in AI foundations, and developing responsible, socially beneficial applications. These priorities span areas including healthcare, education, finance, and Smart Nation (opens in new tab) services. MSRA – Singapore’s research portfolio aligns with these priorities, with a particular focus on building AI systems that are reliable and trustworthy across Southeast Asian contexts. 

The Singapore Economic Development Board (EDB) has been a foundational partner. It supported the lab’s establishment and continues to support joint PhD training through the Industrial Postgraduate Programme (IPP). The relationship has since grown into strategic collaboration. In February 2026, EDB co-hosted an Executive Roundtable on AI in Logistics and Transportation, bringing together more than ten regional industry leaders to explore sectoral challenges.

“Singapore is building deep research capabilities to push the boundaries of AI and other frontier technologies. Microsoft Research Asia – Singapore is an important partner in this effort. Since its launch in 2025, the lab has brought Microsoft’s global research expertise closer to our ecosystem, partnering academia, industry, and government to advance frontier AI research and translate it into real-world applications in areas such as healthcare. MSRA is also developing the next generation of AI research talent who can create meaningful impact for our economy and society.”

— Low Jin Yong, senior vice president and co-head, Digital Industry Singapore

Government engagement extends beyond EDB. In July 2025, IMDA (opens in new tab) co-hosted Microsoft Research Asia Singapore Day (opens in new tab) — Advancing Industrial AI, a full-day event spanning healthcare AI, interactive world simulators, AI agents, and societal AI. MSRA – Singapore has also engaged with the Ministry of Trade and Industry and the Ministry of Health. 

From enduring partnerships to shared discovery

In addition to its deep ties with government agencies, MSRA has built strong connections with Singapore’s academic community, supporting more than 75 research projects with Singapore institutions since 2004. The opening of the Singapore lab has brought new engagements, with nine new projects in its first year with National University of Singapore (NUS) and Nanyang Technological University, Singapore (NTU) across healthcare AI, robotics, AI systems, and multi-agent systems. 

A landmark example is the five-year Framework Research Agreement (opens in new tab) with NUS, announced in 2025, which provides a platform for joint AI research, talent development, and academic exchange. It defines shared research priorities in healthcare, societal AI, spatial intelligence, and data-intensive computing alongside an annual NUS–MSRA Academic Symposium that brings researchers from both organizations together to shape the collaboration agenda. 

“This partnership creates new opportunities for frontier research, talent development, and knowledge exchange, strengthening connections across academia, industry, and the broader innovation ecosystem.”

— Professor Liu Bin, deputy president (Research and Technology), National University of Singapore

The collaboration is already producing concrete research across several frontier areas. NUS Assistant Professor Mike Shou and MSRA’s Xinxing Xu are developing unified multimodal AI for healthcare. NUS assistant professor Yang You and MSRA distinguished scientist Furu Wei are studying systems foundations for diffusion-based language models. NUS assistant professor Jialin Li is working with MSRA on rigorous distributed agreement in multi-agent systems and verifiable machine learning systems.

“Whether during my visit to the Systems Research Group in Beijing or through our ongoing collaborations, I have been impressed by how deeply MSRA researchers engage with difficult problems and how openly they exchange ideas. For young faculty and students, it is a rare environment to explore ambitious ideas.“

— Jialin Li, Sung Kah Kay Assistant Professor, School of Computing, National University of Singapore, MSRA StarTrack Scholar

The NUS–Microsoft Research Asia AI Summer School has reached more than 300 students since its launch in 2025, attracting participants from Singapore, Southeast Asia, and beyond. Through three days of lectures, hands-on workshops, and small-group discussions, the program creates a forum for researchers and students to engage with emerging AI questions and scholars shaping the field. The 2025 program featured a keynote by Doug Burger, technical fellow and corporate vice president of Microsoft, on the AI revolution. Ahead of the 2026 Summer School, Lidong Zhou delivered a lecture on system intelligence, helping frame the conversations that followed. 

“I have seen the NUS–MSRA partnership grow through sustained collaboration in AI research and talent development. What makes this partnership especially meaningful is the opportunity to bring together researchers and students from both communities to learn from one another, pursue new ideas, and create lasting impact together.”

— Tulika Mitra, dean, NUS School of Computing

Beyond NUS, Microsoft Research Asia — Singapore has built connections across Singapore’s research community, including with NTU, Singapore Management University (SMU) (opens in new tab), Agency for Science, Technology and Research (A*STAR), and other institutions. Research arising from these partnerships has already been published at leading venues, including the International Conference on Learning Representations (opens in new tab) (ICLR) and the Conference on Language Modeling (opens in new tab) (COLM). 

From early exposure to research leadership: Cultivating Singapore’s next generation of AI talent

For Microsoft Research Asia, talent development is not a standalone program but a long-term investment in Singapore’s research and innovation capacity, supporting a comprehensive talent pipeline spanning internships, joint PhD supervision, visiting scholar programs, and fellowship initiatives. Through these opportunities, students and early-career researchers gain access to world-class mentorship, a curiosity-driven culture, frontier research challenges, and a global network connecting academia, industry, and real-world applications.

What makes this model distinctive is its continuity. Students may engage with Microsoft Research Asia through workshops or summer schools, gain hands-on experience through its flagship research initiative, Microsoft Research Asia Stars of Tomorrow Internship Program, or pursue doctoral research through joint PhD programs. Some participants later return as faculty collaborators, mentors, or research leaders. Together, these programs form a connected pathway that supports researchers at successive stages of development. 

The IPP represents a sustained form of collaboration, enabling PhD students to conduct long-term research under joint mentorship from academia and industry. Qiming Huang, the first IPP student at MSRA — Singapore, is co-supervised by NUS professor Mike Shou and MSRA researcher Xinxing Xu. Through the program, Huang worked closely with researchers from both institutions, gaining exposure to academic rigor and industry-driven research. At CVPR 2026 (opens in new tab), one of the world’s leading conferences in AI and computer vision, Qiming Huang’s paper on RobotSeg – a foundation model for robot segmentation in image and video – was selected for an oral presentation.  

“Working alongside researchers from both Microsoft Research Asia and NUS changed how I approach AI research. Beyond technical guidance, I learned to frame ambitious questions, collaborate across disciplines, and think about the real-world impact of research

— Qiming Huang, IPP student from National University of Singapore

NTU assistant professor Jianfei Yang, embodies the continuous collaborative opportunities enabled by Microsoft Research Asia. As a student, Yang’s team won the Beauty of Programming competition, an experience that inspired him to pursue a PhD and begin an academic career. Over the years, he watched many of his classmates and close friends grow into outstanding researchers through MSRA programs. Today, he collaborates with MSRA researchers in robotics and jointly supervises students through the IPP, creating for the next generation the same opportunities that inspired his own journey.

“Although I never interned at Microsoft Research Asia myself, it always felt close. Having the opportunity to work alongside a research lab that inspired so many researchers of my generation, and now to co-mentor students together, is both meaningful and rewarding.”

— Jianfei Yang, assistant professor, director of Multimodal embodied AI and Robotic Systems (MARS) Lab, Nanyang Technological University

NTU associate professor Ziwei Liu, himself a former MSRA intern, recently saw one of his students become the only Asia-based recipient of the Microsoft Research global Fellowship award for 2025–2026. 

“As a former intern, I have seen how an environment that brings together world-class researchers, open scientific exchange, and ambitious young talent can shape the next generation of researchers. It is especially meaningful to see my own students now benefit from these opportunities.”

— Ziwei Liu, associate professor, College of Computing and Data Science, Nanyang Technological University

The Stars of Tomorrow Internship Program has engaged more than 90 students from Singapore-based institutions over the years. For many students, these internships serve as a gateway to frontier research and help bridge the gap between academic study and real-world innovation.

For exceptional doctoral students, the Microsoft Research Asia Fellowship provides another pathway for growth, recognition, mentorship, and access to Microsoft’s global research community. Since the program began, 13 students from Singapore-based universities have received the fellowship, many of whom have gone on to become professors, industry research leaders, and technology entrepreneurs around the world.

The next phase: Scaling what works

As Microsoft Research Asia — Singapore enters its second year, the focus is shifting from building foundations to scaling the strongest models of collaboration and demonstrating their impact more broadly.

One year ago, the lab opened with a mission to advance frontier AI in close partnership with the ecosystem around it. That mission is already taking shape through joint research, talent development, and the first partnerships connecting lab work to real-world challenges.

The year ahead is about scaling what works. By deepening these connections, Microsoft Research Asia — Singapore can bring more frontier research into practice and create meaningful value for Singapore, the wider region, industry, and society. 

Opens in a new tab

The post One year in: How Microsoft Research Asia – Singapore is advancing research, partnership and talent for real-world impact appeared first on Microsoft Research.

Categories: Microsoft

Offloaded inference for real-world physical AI robotics

Microsoft Research - Wed, 09/23/2026 - 18:01
At a glance
  • Challenges a core assumption in robotics AI: Our research shows that running physical AI inference exclusively on onboard GPUs can limit robot performance, battery life, and scalability, and that offloading inference to edge or cloud GPUs can offer significant advantages.
  • Demonstrates measurable benefits of inference offloading: Across representative mobile manipulation workloads, offloading improved task success rates, enabled larger AI models, and helped robots respond more effectively in dynamic, real-world environments.
  • Extends robot operating time: Replacing power-hungry onboard AI compute with lightweight onboard hardware and remote inference can substantially improve battery life, enabling robots to operate longer between charges.
  • Introduces a new capability in the Physical AI Toolchain: Developers can now containerize, deploy, and orchestrate robotics AI workloads across robots, edge infrastructure, and the cloud using Kubernetes-based tooling for distributed inference.

Readily-available physical AI, with robotics assisting users in manufacturing, home, and warehouses scenarios, holds immense potential to improve safety, productivity, and assistance across a wide range of tasks. In many ways, AI for the physical world represents a major frontier for AI . Physical AI must operate in open, unpredictable environments, interact with both other robots and people, and work with a diversity of embodiments. Realizing this vision requires advances along three dimensions: robot hardware, embodied AI models, and systems infrastructure for training and inference. While robot hardware and the AI models have advanced rapidly in recent years, we turn our focus on a relatively under-addressed aspect: inference infrastructure of physical AI. Enabling robots to effectively and safely operate in the physical world will require sophisticated systems to handle large volumes of distributed inference compute.

Today, the prevailing approach to physical AI is to provision a GPU onboard the robot, e.g., by wiring a GPU to the robot. In this model, the robot’s inference will be confined to the onboard GPU, and provide the robot with the necessary chunks and sequence of actions for the execution of its tasks. While higher-level planning may be performed in the cloud, task execution typically remains tied to the robot itself. We challenge this assumption.  As physical AI models grow in size and sophistication, the constraints of onboard compute become increasingly apparent. GPUs consume significant power, reduce battery life, add cost and weight, and can limit the ability to run the latest generation of AI models. 

To better understand the systems implications of physical AI, we conducted the first systematic study of robotics workloads. We focused on mobile robotic manipulation, with the canonical task such as “check for rubbish in the kitchen and put it in the trash.” Such a task involves planning the path to the kitchen, perceiving the environment to find rubbish, navigating to the rubbish, picking up the rubbish, and navigating back to the trash can for disposal. We evaluated representative models across three core capabilities: semantic mapping and planning, navigation, and manipulation, as summarized in Figure 2. 

Figure 2: Details of the models used for the different components of mobile manipulation.

Offloading physical AI inference out of the robot improved its response time and accuracy, along with battery lifetime and cost. We evaluated the inference models across a range of onboard, edge, and cloud compute configurations. Details of the specific test hardware are available in our technical report. 

Benefits in task performance: Our evaluation shows offloading inference can significantly improve robot performance across mapping, planning, navigation, and manipulation workloads. Some smaller GPUs could not accommodate the mobile manipulation stack. On GPUs with sufficient memory, mapping and planning slowed by up to 383% compared to an A100, thus limiting the robot’s abilities in dynamic spaces. Navigation showed a 30% drop in its timely detection of obstacles with lighter GPUs. While the VLA models did not dramatically slow down with smaller GPUs, the slowdown was still sufficient to drop their accuracies by 50%. In other words, onboard GPUs limited the performance of the robots while offloading their inference to an on-premise or cloud GPU boosts their operations, as shown in the videos below and quantified in the graphs. As physical AI models continue to grow in size and complexity, the benefits of offloading are likely to become even more pronounced.

Figure 3a: The video shows the handover task with onboard GPUs. Figure 3b: The video shows the handover task when the inference is offloaded. Figure 4: Success rates of robot arms handing over objects to each other when inference is performed with different GPUs (some onboard, and some offloaded). Offloading improves success rates.

Benefits in battery lifetime: Beyond performance, onboard GPUs also significantly drained the battery life of the robot. We compared the increase in battery lifetime by replacing an onboard GPU with a Raspberry Pi-5 board and shipping all the data to the offloaded GPU. The larger onboard GPUs, such as Jetson Thor, drained robot batteries by up to 160% (or a few hours) for even the larger robots.

Figure 5: Impact of offloading GPU inference on the battery life of the robots; the above numbers are for the Stretch-3 robot.

The above results show that offloading GPU inference out of the robot is critical for functioning in the open world with large models and long battery lifetimes. Nonetheless, offloading inference out of the robot involves a complex tradeoff involving performance, network latency and bandwidth, and available GPU resources. We believe that our measurement study will inform the design of physical AI inference systems.

video series

On Second Thought

A video series with Sinead Bovell built around the questions everyone’s asking about AI. With expert voices from across Microsoft, we break down the tension and promise of this rapidly changing technology, exploring what’s evolving and what’s possible.

Explore the series Opens in a new tab Toolset for automatic offload

We have built a toolset for easy inference offloading out of the robot and distributing inference between the edge GPU and cloud. Kubernetes is a natural platform to provide a uniform abstraction to distribute robotic AI between the robot’s compute, edge GPU, and overflowing to the cloud. The toolset allows automatic containerization and offloading of robotics workloads using declarative specifications, distributes physical AI containers with smart policies using Kubernetes, and integrates with robotic simulators, LeRobot, and ROS2 for easy development. The sequence of steps below shows how the toolset can be prompted with what to offload, and how it creates a separate container for GPU inference and offloads the same.

Figure 6: Steps in the offloading toolset with containerization and deployment.

Microsoft has recently released the Physical AI Toolchain (opens in new tab) for operationalizing physical intelligence at scale. Physical AI Toolchain is an open-source, production-ready framework that integrates Microsoft Azure (opens in new tab) cloud services with NVIDIA’s (opens in new tab) physical AI stack, accelerating robotics and physical AI developers to automate and scale data curation, augmentation, and evaluation across perception, mobility, imitation learning, and reinforcement learning pipelines. We are announcing the addition of an industry-first capability for offloaded physical AI inference for robots as part of the Physical AI Toolchain. This release includes example projects for offloading inference of a SO-101 and a UR10e. The videos below show the offloading of the inference of Microsoft’s Rho model, targeted at dual-arm robots, to a Jetson Thor GPU, which controls the actions of the Mobile Aloha robot (opens in new tab).  

Figure 7a: Offloading of GPU inference of Microsoft’s Rho model controlling the Mobile Aloha robot for interactions with the BusyBox to press the blue button. Figure 7b: Offloading of GPU inference of Microsoft’s Rho model controlling the Mobile Aloha robot for interactions with the BusyBox to turn the knob to position 4.

Check out the inference offload feature, look into the source code, and let us know your feedback. We have already tested it with many real-world use cases, and look forward to hearing about your deployment experiences.

Opens in a new tab

The post Offloaded inference for real-world physical AI robotics appeared first on Microsoft Research.

Categories: Microsoft

Improving synthesis prediction of small molecules at scale with RetroChimera

Microsoft Research - Mon, 09/21/2026 - 17:30
At a glance
  • We report on the recent publication of our retrosynthesis model RetroChimera in the journal Nature (opens in new tab).
  • The paper describes the model’s architecture as well as extensive validation studies, including the model’s ability to recall rare reaction types, and successful zero-shot transfer and fine-tuning on proprietary datasets.
  • We open-source RetroChimera’s implementation and weights in the hope that it will enable researchers to accelerate development of new medicinally relevant molecules and advanced materials. 

Developing new medicines and materials requires making new molecules but planning how to make them is still largely manual, time-consuming, and costly. RetroChimera automatically proposes high-quality synthesis routes. The model combines two strong models with complementary strengths, learning how to rank their proposals to produce better predictions than either alone. In blind tests, PhD-level chemists prefer RetroChimera’s individual reaction predictions over preceding models and recorded literature reactions. 

Custom-made molecules are unlocking advances in modern medicine, smart materials, and sustainable agriculture. Yet, progress is slowed by chemical synthesis—the time-consuming process of making new molecules from simpler building blocks in the lab. In addition, synthesis is a significant driver of drug development costs. So even as computational methods make it possible to explore large numbers of novel molecules, finding practical ways to synthesize them remains a critical challenge. 

Figure 1: Planning a synthesis by working backward. Retrosynthesis starts with a target molecule and proposes successive disconnections into simpler precursors until purchasable building blocks are reached. The highlighted path shows a complete synthesis route; pale branches illustrate alternatives explored along the way. Circles represent molecules and squares represent reactions. For clarity, only a few branches are illustrated, with chemical structures shown for the target, one intermediate, and selected building blocks. 

Retrosynthesis approaches this problem by working backwards from a target molecule, breaking it down step by step into simpler precursors (Figure 1). This process is comparable to playing strategic board games like chess and Go. It involves contemplating a wide range of possible immediate moves, or individual disconnections, while also requiring high-level strategic thinking to reach the end-to-end synthesis plan. However, the number of possible moves in retrosynthesis is much larger than in board games, and it is not obvious which moves would be available for a given molecule. Existing systems face major challenges, including recalling rare but strategically important reactions, robustness beyond the training distribution, and aligning with chemists’ expectations. As a result, retrosynthesis often requires highly specialized expertise, which hinders scaling and automation of scientific discovery.

Figure 2: Our framework for ensemble-based retrosynthesis with learned re-ranking which underpins RetroChimera. The ensemble receives a target molecule as the input, which is then processed by the sub-models. The model outputs are aggregated using a learning-to-rank strategy. While in this work we only investigate deep learning models as prediction sources (solid boxes), it is possible to add additional sources, for example calls to reaction databases or human-in-the-loop queries (dashed box). 

In a paper recently published in the journal Nature (opens in new tab), we present RetroChimera (opens in new tab), a new framework for retrosynthesis prediction. It is built around two models (Figure 2). R-SMILES 2, a Transformer-based de-novo model, predicts precursor molecules directly from the input molecule. This gives it the flexibility to learn reaction patterns directly from data. However, its unconstrained generation can also make it prone to hallucination.

NeuralLoc, in contrast, is a graph neural network- (GNN) based model that encodes both the target molecule and reaction templates as graphs. It selects reaction templates and predicts where they should be applied to the target molecule. Its predictions are grounded in reaction patterns extracted from the training data, so it tends to produce more accurate and reliable outputs. But it’s more constrained when encountering reactions not covered by the template library. 

These differences actually turn out to be a strength. Rather than making the same kinds of predictions, the two models capture complementary patterns in chemistry and specialize in different reaction types. R-SMILES 2 performs particularly well on reactions that involve large changes over the course of the reaction, while NeuralLoc excels in reactions of low precedence and those involving more localized changes.

RetroChimera combines the ranked predictions of both sub-models using a learned ensembling strategy. Each model assigns a learned, rank-dependent vote to each predicted reactant set, and votes are added when both models propose the same reaction. By learning how much to trust each model at different ranks, RetroChimera can leverage their complementary strengths, approximately matching the better-performing sub-model across reaction classes.

Figure 3: Expert assessment of multistep synthesis routes. Left: Ratings of individual reaction steps. Right: Complete routes accepted or rejected for ten challenging targets. RetroChimera succeeded on nine targets, versus five for the de novo model, four for the editing model, and two for NeuralSym, a strong baseline model. 

As a result, RetroChimera performs strongly across both common and rare reaction classes and produces retrosynthesis predictions that better align with chemists’ judgment (Figure 3). In blind tests, expert chemists preferred disconnections of complex molecules suggested by RetroChimera over those obtained from its constituent sub-models, as well as those from more established approaches, and even from the test set itself.

We believe RetroChimera could help researchers identify promising synthesis route more efficiently, supporting faster design-make-test cycle across molecular science applications, including drug discovery and design of smart materials. RetroChimera could enable chemists to assess more—and more complex—candidate molecules at large scale. Paired with increasing levels of laboratory automation, we expect further acceleration toward closed-loop, self-improving systems for synthesis planning and execution.  

RetroChimera is available on GitHub (opens in new tab) (MIT license) and accessible via Microsoft Foundry (opens in new tab). For instructions on how to access the checkpoint, we refer to the GitHub repository.

We invite the broader chemistry community to experiment with RetroChimera, helping us identify its strengths and shortcomings so we can enhance it in the future. We are looking forward to hearing how it performs on various targets you care about!

For a deeper look at the findings, including evaluation experiments with our external research collaborators, see the full Nature publication (opens in new tab) and the accompanying Microsoft Source article (opens in new tab). 

Publication OSS repository Microsoft Foundry Opens in a new tab

The post Improving synthesis prediction of small molecules at scale with RetroChimera appeared first on Microsoft Research.

Categories: Microsoft

GigaPath-Flash and GigaTIME-Flash: Toward population-scale discovery with efficient pathology foundation models

Microsoft Research - Mon, 08/31/2026 - 18:00
At a glance
  • The Flash family extends GigaPath and GigaTIME with dramatically improved efficiency, making large-scale pathology research more accessible and practical.
  • A distilled pathology foundation model backbone reduces computational requirements without sacrificing performance, enabling repeated analyses across larger patient cohorts.
  • These open models support population-scale discovery, helping researchers investigate disease biology, biomarkers, and clinical outcomes across diverse cancer datasets.

GigaPath (opens in new tab) and GigaTIME (opens in new tab) demonstrated how foundation models can support whole-slide analysis and tumor microenvironment modeling from routinely collected pathology data. GigaPath-Flash and GigaTIME-Flash make these capabilities substantially more efficient, enabling researchers to analyze larger cohorts, run more experiments, and move toward population-scale discovery. GigaPath-Flash and GigaTIME-Flash are research models. They are not intended or validated for clinical use, including diagnosis, prognosis, treatment selection, or other patient-care decisions. Performance may vary across datasets, scanners, institutions, populations, and use cases.

The scale opportunity in computational pathology

Histopathology is among the richest and most widely available sources of information in cancer research. Every tissue biopsy produces a whole-slide image that captures cellular morphology at subcellular resolution — and hospitals generate millions of these slides each year. This data contains information relevant to diagnosis, prognosis, treatment selection, and the biology of the tumor microenvironment.

Foundation models have begun to unlock this information at scale. But whole-slide images are large — often exceeding a gigapixel — and applying a foundation model to even a single slide requires processing thousands of image tiles. When a research question involves tens of thousands of patients, the computational cost grows quickly. And population-scale discovery is not a single model run: it requires repeated cycles of feature extraction, statistical analysis, hypothesis testing, and validation across patient subgroups, biomarkers, and clinical endpoints. 

Computational cost limits the number of patients, datasets, tasks, and hypotheses that researchers can study. To realize the full potential of pathology foundation models, we need models that can be applied repeatedly and affordably across large patient populations.

From GigaPath and GigaTIME to the -Flash family

GigaPath (Nature, 2024) (opens in new tab) is a whole-slide foundation model pretrained on large-scale real-world histopathology data from Providence. Unlike models that operate only at the tile level, GigaPath learns contextualized representations of entire slides, capturing both local cellular patterns and global tissue architecture.

GigaTIME (Cell, 2026) (opens in new tab) extends this line of work to tumor microenvironments. Trained on 40 million cells with paired H&E and multiplex immunofluorescence (mIF) data, GigaTIME translates routine H&E images into virtual spatial proteomics maps across 21 protein channels. Applied to over 14,000 cancer patients, it generated a virtual population that uncovered more than 1,200 statistically significant associations between immune cell states and clinical biomarkers.

GigaPath and GigaTIME addressed the scale of pathology data and biological discovery. And now, the Flash family of models addresses scale of experimentation.

Figure 1: Overview of the GigaPath/GigaTIME model family. GigaPath-Flash provides efficient tile and slide encoders at 22M parameters. GigaTIME-Flash predicts spatial proteomics from H&E, replacing the CNN backbone with the distilled ViT-S encoder. Introducing GigaPath-Flash and GigaTIME-Flash

The Flash family shares a core design goal: preserving useful pathology representations while substantially reducing the computational resources required to generate and use them. Both models are built on a common efficient backbone — a compact ViT-S tile encoder distilled from the original billion-parameter GigaPath encoder — and both are released under the Apache 2.0 license. 

GigaPath-Flash

GigaPath-Flash is an efficient foundation model for whole-slide representation learning. It combines a 22M-parameter ViT-S tile encoder with a 21M-parameter LongNet slide encoder. The tile encoder is distilled from the original GigaPath ViT-g teacher, transferring the representational capacity of a billion-parameter model into a backbone that is an order of magnitude smaller. The slide encoder contextualizes all tile embeddings via dilated attention, scaling linearly with the number of tiles.

On slide-level classification benchmarks (PANDA prostate grading and EBRAINS brain tumor subtyping), GigaPath-Flash achieves the lowest inference cost among whole-slide pretrained models while retaining competitive performance — scoring within 3% of the original GigaPath at roughly 50 times less compute.

Figure 2: Efficiency–performance trade-off on whole-slide benchmarks. GigaPath-Flash (red, top-left) achieves competitive performance at substantially lower computational cost than other whole-slide pretrained models. GigaTIME-Flash

GigaTIME-Flash replaces the CNN backbone of the original GigaTIME with the GigaPath-Flash ViT-S encoder, paired with a lightweight convolutional decoder for H&E-to-mIF translation. The model is fine-tuned using LoRA adapters, keeping the pretrained encoder weights largely frozen.

On both in-distribution and out-of-distribution cohorts spanning brain, breast, colon, and lung cancers, GigaTIME-Flash matches or improves upon the original GigaTIME in spatial protein prediction quality. The gains are particularly notable on out-of-distribution data, suggesting that the foundation model backbone improves generalization to previously unseen tissue types.

Figure 3: Mean windowed Pearson correlation for GigaTIME and GigaTIME-Flash on the GigaTIME test set and four out-of-distribution Prov-TMA cohorts. GigaTIME-Flash matches or improves upon the original across all cohorts. Efficiency without giving up the foundation

The efficiency gains of the Flash models are substantial:

ModelTypeEfficiency gainGigaPath-FlashWhole-slide Foundation Model~50× less compute, 97% of predictive performance compared to GigaPathGigaTIME-FlashSpatial Proteomics~6× faster, ~8× less memory, better predictive performance compared to GigaTIME

For a single slide, these differences reduce runtime and hardware requirements. Across tens of thousands of slides, they can determine whether an experiment is practical at all. To illustrate, we estimate the wall-clock time for generating virtual mIF across cohorts of different sizes on a single A100 GPU, assuming approximately 10,000 tiles per slide:

Model1,000 slides100K slides1M slidesGigaTIME-Flash~2 GPU-hours~7 GPU-days~70 GPU-daysGigaTIME~7 GPU-hours~30 GPU-days~300 GPU-daysEstimates assume ~10,000 tiles per slide, batch size 128, single NVIDIA A100 GPU. Actual runtime depends on slide size, tiling resolution, and hardware. Figure 4: GigaTIME efficiency scaling. Left: throughput (tiles/sec) vs. batch size. Right: peak GPU memory (GB) vs. batch size. GigaTIME-Flash scales to over 1,600 tiles/sec while using a fraction of the memory. An open model release

Both GigaPath-Flash and GigaTIME-Flash are released as open-weight models under the Apache 2.0 license. Model weights and code are available on HuggingFace:

This is an early research release. Our current evaluations cover a limited set of benchmarks and cohorts, and broader validation across tasks, scanners, and patient populations is still needed. We expect the most valuable applications of these models to include scientific questions, cohorts, and use cases beyond those in our initial experiments. We welcome community evaluation, and we are equally interested in reports of where these models work well and where they fall short.

Downstream clinical applications will require additional multi-institutional and prospective validation.

ACADEMIC CONFERENCE

Microsoft at SOSP 2026

Academic and industrial participants present research and experience papers that cover the full range of theory and practice of computer systems software.

Learn more Opens in a new tab Efficiency as an enabler of discovery

GigaPath and GigaTIME demonstrated what pathology foundation models can learn from whole slides and tumor tissue. GigaPath-Flash and GigaTIME-Flash are a step toward making those capabilities usable across larger populations, more experiments, and a broader research community.

By making pathology foundation models more efficient, we hope to expand the scale of the scientific questions researchers can ask.

Acknowledgements

GigaPath-Flash and GigaTIME-Flash are joint work across Microsoft Research, the University of Washington, and Providence. For technical details, see the paper.

Paper co-authors: Naoto Usuyama, Jeya Maria Jose Valanarasu, Sicong Yao, Hanwen Xu, Jaspreet Bagga, Guanghui Qin, Robert E. Kramer, Cliff Wong, Soohee Lee, Hao Qiu, Theodore Zhengde Zhao, Racheli Ben Shimol, Angela Crabtree, Kevin Matlock, Eduardo Alejandro Lozano Garcia, Naiteek Sangani, Alberto Santamaria-Pang, Maximilian Rokuss, Yashna Hasija, Naisargi Manishkumar Patel, Jason Entenmann, Alexandra Q. Bartlett, Bill J. Wright, Bernard A. Fox, Brian Piening, Sheng Zhang, Sheng Wang, Tristan Naumann, Carlo Bifulco, Hoifung Poon

Opens in a new tab

The post GigaPath-Flash and GigaTIME-Flash: Toward population-scale discovery with efficient pathology foundation models appeared first on Microsoft Research.

Categories: Microsoft

GigaPath-Flash and GigaTIME-Flash: Toward population-scale discovery with efficient pathology foundation models

Microsoft Research - Mon, 08/31/2026 - 18:00
At a glance
  • The Flash family extends GigaPath and GigaTIME with dramatically improved efficiency, making large-scale pathology research more accessible and practical.
  • A distilled pathology foundation model backbone reduces computational requirements without sacrificing performance, enabling repeated analyses across larger patient cohorts.
  • These open models support population-scale discovery, helping researchers investigate disease biology, biomarkers, and clinical outcomes across diverse cancer datasets.

GigaPath (opens in new tab) and GigaTIME (opens in new tab) demonstrated how foundation models can support whole-slide analysis and tumor microenvironment modeling from routinely collected pathology data. GigaPath-Flash and GigaTIME-Flash make these capabilities substantially more efficient, enabling researchers to analyze larger cohorts, run more experiments, and move toward population-scale discovery. GigaPath-Flash and GigaTIME-Flash are research models. They are not intended or validated for clinical use, including diagnosis, prognosis, treatment selection, or other patient-care decisions. Performance may vary across datasets, scanners, institutions, populations, and use cases.

The scale opportunity in computational pathology

Histopathology is among the richest and most widely available sources of information in cancer research. Every tissue biopsy produces a whole-slide image that captures cellular morphology at subcellular resolution — and hospitals generate millions of these slides each year. This data contains information relevant to diagnosis, prognosis, treatment selection, and the biology of the tumor microenvironment.

Foundation models have begun to unlock this information at scale. But whole-slide images are large — often exceeding a gigapixel — and applying a foundation model to even a single slide requires processing thousands of image tiles. When a research question involves tens of thousands of patients, the computational cost grows quickly. And population-scale discovery is not a single model run: it requires repeated cycles of feature extraction, statistical analysis, hypothesis testing, and validation across patient subgroups, biomarkers, and clinical endpoints. 

Computational cost limits the number of patients, datasets, tasks, and hypotheses that researchers can study. To realize the full potential of pathology foundation models, we need models that can be applied repeatedly and affordably across large patient populations.

From GigaPath and GigaTIME to the -Flash family

GigaPath (Nature, 2024) (opens in new tab) is a whole-slide foundation model pretrained on large-scale real-world histopathology data from Providence. Unlike models that operate only at the tile level, GigaPath learns contextualized representations of entire slides, capturing both local cellular patterns and global tissue architecture.

GigaTIME (Cell, 2026) (opens in new tab) extends this line of work to tumor microenvironments. Trained on 40 million cells with paired H&E and multiplex immunofluorescence (mIF) data, GigaTIME translates routine H&E images into virtual spatial proteomics maps across 21 protein channels. Applied to over 14,000 cancer patients, it generated a virtual population that uncovered more than 1,200 statistically significant associations between immune cell states and clinical biomarkers.

GigaPath and GigaTIME addressed the scale of pathology data and biological discovery. And now, the Flash family of models addresses scale of experimentation.

Figure 1: Overview of the GigaPath/GigaTIME model family. GigaPath-Flash provides efficient tile and slide encoders at 22M parameters. GigaTIME-Flash predicts spatial proteomics from H&E, replacing the CNN backbone with the distilled ViT-S encoder. Introducing GigaPath-Flash and GigaTIME-Flash

The Flash family shares a core design goal: preserving useful pathology representations while substantially reducing the computational resources required to generate and use them. Both models are built on a common efficient backbone — a compact ViT-S tile encoder distilled from the original billion-parameter GigaPath encoder — and both are released under the Apache 2.0 license. 

GigaPath-Flash

GigaPath-Flash is an efficient foundation model for whole-slide representation learning. It combines a 22M-parameter ViT-S tile encoder with a 21M-parameter LongNet slide encoder. The tile encoder is distilled from the original GigaPath ViT-g teacher, transferring the representational capacity of a billion-parameter model into a backbone that is an order of magnitude smaller. The slide encoder contextualizes all tile embeddings via dilated attention, scaling linearly with the number of tiles.

On slide-level classification benchmarks (PANDA prostate grading and EBRAINS brain tumor subtyping), GigaPath-Flash achieves the lowest inference cost among whole-slide pretrained models while retaining competitive performance — scoring within 3% of the original GigaPath at roughly 50 times less compute.

Figure 2: Efficiency–performance trade-off on whole-slide benchmarks. GigaPath-Flash (red, top-left) achieves competitive performance at substantially lower computational cost than other whole-slide pretrained models. GigaTIME-Flash

GigaTIME-Flash replaces the CNN backbone of the original GigaTIME with the GigaPath-Flash ViT-S encoder, paired with a lightweight convolutional decoder for H&E-to-mIF translation. The model is fine-tuned using LoRA adapters, keeping the pretrained encoder weights largely frozen.

On both in-distribution and out-of-distribution cohorts spanning brain, breast, colon, and lung cancers, GigaTIME-Flash matches or improves upon the original GigaTIME in spatial protein prediction quality. The gains are particularly notable on out-of-distribution data, suggesting that the foundation model backbone improves generalization to previously unseen tissue types.

Figure 3: Mean windowed Pearson correlation for GigaTIME and GigaTIME-Flash on the GigaTIME test set and four out-of-distribution Prov-TMA cohorts. GigaTIME-Flash matches or improves upon the original across all cohorts. Efficiency without giving up the foundation

The efficiency gains of the Flash models are substantial:

ModelTypeEfficiency gainGigaPath-FlashWhole-slide Foundation Model~50× less compute, 97% of predictive performance compared to GigaPathGigaTIME-FlashSpatial Proteomics~6× faster, ~8× less memory, better predictive performance compared to GigaTIME

For a single slide, these differences reduce runtime and hardware requirements. Across tens of thousands of slides, they can determine whether an experiment is practical at all. To illustrate, we estimate the wall-clock time for generating virtual mIF across cohorts of different sizes on a single A100 GPU, assuming approximately 10,000 tiles per slide:

Model1,000 slides100K slides1M slidesGigaTIME-Flash~2 GPU-hours~7 GPU-days~70 GPU-daysGigaTIME~7 GPU-hours~30 GPU-days~300 GPU-daysEstimates assume ~10,000 tiles per slide, batch size 128, single NVIDIA A100 GPU. Actual runtime depends on slide size, tiling resolution, and hardware. Figure 4: GigaTIME efficiency scaling. Left: throughput (tiles/sec) vs. batch size. Right: peak GPU memory (GB) vs. batch size. GigaTIME-Flash scales to over 1,600 tiles/sec while using a fraction of the memory. An open model release

Both GigaPath-Flash and GigaTIME-Flash are released as open-weight models under the Apache 2.0 license. Model weights and code are available on HuggingFace:

This is an early research release. Our current evaluations cover a limited set of benchmarks and cohorts, and broader validation across tasks, scanners, and patient populations is still needed. We expect the most valuable applications of these models to include scientific questions, cohorts, and use cases beyond those in our initial experiments. We welcome community evaluation, and we are equally interested in reports of where these models work well and where they fall short.

Downstream clinical applications will require additional multi-institutional and prospective validation.

Spotlight: Microsoft research newsletter

Microsoft Research Newsletter

Stay connected to the research community at Microsoft.

Subscribe today Opens in a new tab Efficiency as an enabler of discovery

GigaPath and GigaTIME demonstrated what pathology foundation models can learn from whole slides and tumor tissue. GigaPath-Flash and GigaTIME-Flash are a step toward making those capabilities usable across larger populations, more experiments, and a broader research community.

By making pathology foundation models more efficient, we hope to expand the scale of the scientific questions researchers can ask.

Acknowledgements

GigaPath-Flash and GigaTIME-Flash are joint work across Microsoft Research, the University of Washington, and Providence. For technical details, see the paper.

Paper co-authors: Naoto Usuyama, Jeya Maria Jose Valanarasu, Sicong Yao, Hanwen Xu, Jaspreet Bagga, Guanghui Qin, Robert E. Kramer, Cliff Wong, Soohee Lee, Hao Qiu, Theodore Zhengde Zhao, Racheli Ben Shimol, Angela Crabtree, Kevin Matlock, Eduardo Alejandro Lozano Garcia, Naiteek Sangani, Alberto Santamaria-Pang, Maximilian Rokuss, Yashna Hasija, Naisargi Manishkumar Patel, Jason Entenmann, Alexandra Q. Bartlett, Bill J. Wright, Bernard A. Fox, Brian Piening, Sheng Zhang, Sheng Wang, Tristan Naumann, Carlo Bifulco, Hoifung Poon

Opens in a new tab

The post GigaPath-Flash and GigaTIME-Flash: Toward population-scale discovery with efficient pathology foundation models appeared first on Microsoft Research.

Categories: Microsoft

MindTopo reveals VLMs’ spatial reasoning abilities

Microsoft Research - Wed, 08/12/2026 - 18:00
At a glance
  • MindTopo is a new benchmark for testing topological reasoning in AI, evaluating whether multimodal models can understand concepts such as connectivity, enclosure, order, separation, and knots.
  • The benchmark measures both reasoning and planning, testing not only whether models can recognize topological relationships in static images but also whether they can preserve and manipulate those relationships through a sequence of actions.
  • Current multimodal models perform much better on static recognition than interactive tasks, suggesting they struggle to maintain a consistent understanding of topology over time.
  • Failures often emerge during planning rather than perception, with models losing track of structural relationships as scenes change or proposing actions that violate physical constraints.
  • The findings highlight an important opportunity to advance AI systems for robotics and interactive environments, where understanding what stays connected, enclosed, ordered, or knotted is essential for reliable decision-making.

Can AI determine whether two rooms remain connected after a wall is added? Can it recognize whether an animal is inside a fence, distinguish a true knot from a tangled loop, or rearrange several ropes without allowing them to pass through one another?

These questions concern 3D topology, a form of spatial understanding based not on exact distances, angles, or shapes, but on structural relationships that persist as objects bend, stretch, or deform. Connectivity, enclosure, ordering, and knottedness are examples of topological properties. These properties are a foundational layer of human spatial understanding in Cognitive Science, yet they remain largely absent from how multimodal AI systems are evaluated.

In a new research study, we introduce MindTopo (opens in new tab), a benchmark designed to evaluate whether multimodal large language models possess this kind of topological intuition. Our findings reveal a substantial gap between recognizing topology in a static image and maintaining an innate understanding of that topology while planning and acting. Current models can sometimes identify a connected path, enclosed region, or knot in a single scene, but that understanding often breaks down once the model must manipulate the scene through a sequence of actions.

How MindTopo defines topological space

Most spatial evaluations for multimodal models focus on Euclidean properties such as distance, direction, size, and relative position. Inspired by Piaget and other cognitive literature’s classification of topological ability, MindTopo organizes its tasks around the following five categories:

  • Continuity asks whether a path or object remains unbroken.
  • Separation asks whether nearby elements form one structure or distinct parts.
  • Order tracks how elements are arranged along a path or through a transformation.
  • Enclosure tests whether a boundary creates an inside and an outside.
  • Knots tests whether ropes are truly knotted or linked rather than merely tangled in appearance.

Each category is evaluated at two cognitive levels. In reasoning tasks, a model examines one or more rendered scenes and answers a question about their topological structure: whether two points in a maze are connected, whether the sheep are inside the fence, whether a rope is truly knotted. In planning tasks, the model interacts with a simulated environment and selects actions that must create, preserve, or remove a particular relation, such as rotating pipe segments, drawing a separating path, rearranging blocks, trapping a moving agent, or untangling ropes. The environments enforce legal actions, so a model cannot solve a rope puzzle by passing one strand through another.

Figure 1. MindTopo pairs questions about static scenes with interactive tasks that require models to preserve or change the same topological relations. 

All scenes are generated from controlled simulators, which provide exact ground truth and adjustable difficulty. That control makes it possible to separate two failure modes that otherwise look alike: a model that fails because a scene is visually complex, and a model that fails because it cannot maintain the underlying relationship as objects move.

Figure 2. MindTopo maps reasoning and planning tasks to continuity, separation, order, enclosure, and knots.  Seeing topology is not the same as acting on it

Across a broad set of proprietary and open-weight models, performance was consistently stronger on static reasoning than on interactive planning, and both remained well below human performance. The contrast was especially clear when success depended on preserving a relationship across many actions.

The error patterns help locate the problem. Static mistakes usually began with perception, such as missing a wall, opening, or crossing. Planning mistakes appeared after the scene had been understood. Models followed a locally plausible move without tracking its later consequences, lost the task over multiple turns, or proposed an action that violated the environment’s dynamics.

Microsoft Research at BUILD 2026

Giving developers a hands-on look at some of the many AI-based models and tools they can use to accelerate innovation, enhance their capabilities, and quickly transform ideas into prototypes.

Learn more Opens in a new tab What generative tools reveal

We also tested whether image and video generation could help models maintain an understanding of topological relationships. Image generation sometimes helped when the relevant relation was visible in a single frame, but it remained unreliable across a sequence of crossings or moves. Video rollouts frequently altered topology or violated task dynamics. Visual simulation appeared useful only to the extent that it preserved structural constraints over time.

Building agents that preserve structure

MindTopo is intended as a controlled diagnostic for this gap. Robots, accessibility tools, and interactive assistants must understand not only where objects are, but also what remains connected, enclosed, ordered, or knotted as actions unfold. Closing that gap may require models that carry an explicit topological state, or world models whose predictions preserve topology by construction.

Opens in a new tab

The post MindTopo reveals VLMs’ spatial reasoning abilities appeared first on Microsoft Research.

Categories: Microsoft

MindTopo reveals VLMs’ spatial reasoning abilities

Microsoft Research - Wed, 08/12/2026 - 18:00
At a glance
  • MindTopo is a new benchmark for testing topological reasoning in AI, evaluating whether multimodal models can understand concepts such as connectivity, enclosure, order, separation, and knots.
  • The benchmark measures both reasoning and planning, testing not only whether models can recognize topological relationships in static images but also whether they can preserve and manipulate those relationships through a sequence of actions.
  • Current multimodal models perform much better on static recognition than interactive tasks, suggesting they struggle to maintain a consistent understanding of topology over time.
  • Failures often emerge during planning rather than perception, with models losing track of structural relationships as scenes change or proposing actions that violate physical constraints.
  • The findings highlight an important opportunity to advance AI systems for robotics and interactive environments, where understanding what stays connected, enclosed, ordered, or knotted is essential for reliable decision-making.

Can AI determine whether two rooms remain connected after a wall is added? Can it recognize whether an animal is inside a fence, distinguish a true knot from a tangled loop, or rearrange several ropes without allowing them to pass through one another?

These questions concern 3D topology, a form of spatial understanding based not on exact distances, angles, or shapes, but on structural relationships that persist as objects bend, stretch, or deform. Connectivity, enclosure, ordering, and knottedness are examples of topological properties. These properties are a foundational layer of human spatial understanding in Cognitive Science, yet they remain largely absent from how multimodal AI systems are evaluated.

In a new research study, we introduce MindTopo (opens in new tab), a benchmark designed to evaluate whether multimodal large language models possess this kind of topological intuition. Our findings reveal a substantial gap between recognizing topology in a static image and maintaining an innate understanding of that topology while planning and acting. Current models can sometimes identify a connected path, enclosed region, or knot in a single scene, but that understanding often breaks down once the model must manipulate the scene through a sequence of actions.

How MindTopo defines topological space

Most spatial evaluations for multimodal models focus on Euclidean properties such as distance, direction, size, and relative position. Inspired by Piaget and other cognitive literature’s classification of topological ability, MindTopo organizes its tasks around the following five categories:

  • Continuity asks whether a path or object remains unbroken.
  • Separation asks whether nearby elements form one structure or distinct parts.
  • Order tracks how elements are arranged along a path or through a transformation.
  • Enclosure tests whether a boundary creates an inside and an outside.
  • Knots tests whether ropes are truly knotted or linked rather than merely tangled in appearance.

Each category is evaluated at two cognitive levels. In reasoning tasks, a model examines one or more rendered scenes and answers a question about their topological structure: whether two points in a maze are connected, whether the sheep are inside the fence, whether a rope is truly knotted. In planning tasks, the model interacts with a simulated environment and selects actions that must create, preserve, or remove a particular relation, such as rotating pipe segments, drawing a separating path, rearranging blocks, trapping a moving agent, or untangling ropes. The environments enforce legal actions, so a model cannot solve a rope puzzle by passing one strand through another.

Figure 1. MindTopo pairs questions about static scenes with interactive tasks that require models to preserve or change the same topological relations. 

All scenes are generated from controlled simulators, which provide exact ground truth and adjustable difficulty. That control makes it possible to separate two failure modes that otherwise look alike: a model that fails because a scene is visually complex, and a model that fails because it cannot maintain the underlying relationship as objects move.

Figure 2. MindTopo maps reasoning and planning tasks to continuity, separation, order, enclosure, and knots.  Seeing topology is not the same as acting on it

Across a broad set of proprietary and open-weight models, performance was consistently stronger on static reasoning than on interactive planning, and both remained well below human performance. The contrast was especially clear when success depended on preserving a relationship across many actions.

The error patterns help locate the problem. Static mistakes usually began with perception, such as missing a wall, opening, or crossing. Planning mistakes appeared after the scene had been understood. Models followed a locally plausible move without tracking its later consequences, lost the task over multiple turns, or proposed an action that violated the environment’s dynamics.

Spotlight: Event Series

Microsoft Research Forum

Join us for a continuous exchange of ideas about research in the era of general AI. Watch the latest episodes on demand.

Watch on-demand Opens in a new tab What generative tools reveal

We also tested whether image and video generation could help models maintain an understanding of topological relationships. Image generation sometimes helped when the relevant relation was visible in a single frame, but it remained unreliable across a sequence of crossings or moves. Video rollouts frequently altered topology or violated task dynamics. Visual simulation appeared useful only to the extent that it preserved structural constraints over time.

Building agents that preserve structure

MindTopo is intended as a controlled diagnostic for this gap. Robots, accessibility tools, and interactive assistants must understand not only where objects are, but also what remains connected, enclosed, ordered, or knotted as actions unfold. Closing that gap may require models that carry an explicit topological state, or world models whose predictions preserve topology by construction.

Opens in a new tab

The post MindTopo reveals VLMs’ spatial reasoning abilities appeared first on Microsoft Research.

Categories: Microsoft

Introducing CARE-X: Towards Clinically Useful Radiology VLMs with Auxiliary Supervision, Reward-Aligned Learning, and Tool-Augmented Measurement

Microsoft Research - Tue, 08/11/2026 - 18:00

Research Note: CARE-X is a research model and not a Microsoft product offering or medical device. It has not been cleared or approved by any regulatory authority and is not intended for clinical diagnosis, screening, or patient care. The results described below are retrospective research findings and do not establish the safety, effectiveness, or suitability of CARE-X for any clinical use. References to potential workflows describe areas for future research, not currently available capabilities or recommended uses. 

At a glance
  • The challenge: Chest X-ray interpretation spans diverse tasks that require both expressive report generation and calibrated diagnostic predictions.
  • CARE-X is a unified chest X-ray VLM for diverse clinical interpretation tasks. It combines generation and structured prediction to provide both free-text reasoning and deterministic outputs.
  • CARE-X uses reinforcement learning (DAPO) to reward clinical correctness in a multi-task setting.
  • In a separate research experiment from CARE-X, we paired Qwen3-VL-4B-Instruct with deterministic measurement tools to evaluate whether direct computation could improve performance on measurement-dependent conditions compared with visual approximation alone. 
  • Validated on real-world Indian clinical data from Narayana Health, including rare ICU pathologies and CT-confirmed enlargement conditions.
What radiologists need: Task diversity, flexibility, and clinical fidelity

A clinically useful radiology AI system must support a wide range of tasks, adapt to different workflows, and produce outputs that are medically accurate.

Radiologists and other clinicians use chest X-rays for many different purposes. A clinically useful AI system must be able to support that range of tasks. It may be asked to generate detailed findings and concise impressions for a report, answer questions about the presence, absence, or location of a finding, identify medical devices and assess their placement, or pinpoint exactly where an abnormality appears in an image.

These tasks also require different kinds of outputs, from narrative reports to calibrated diagnostic scores. And above all, they require clinical accuracy. A report could ostensibly be perfectly written yet clinically wrong if it misses a finding, reverses a negation, or misidentifies a location. Certain findings could be trivial in one context and vital to identify in another.

CARE-X was developed as a research model to explore how a unified approach can address these diverse demands. The system combines generative and discriminative capabilities, clinically aligned optimization, and tool-based reasoning to support a broader range of radiology workflows while maintaining clinical fidelity.

Microsoft Research at BUILD 2026

Giving developers a hands-on look at some of the many AI-based models and tools they can use to accelerate innovation, enhance their capabilities, and quickly transform ideas into prototypes.

Learn more Opens in a new tab Gaps in current radiology vision-language models

Despite the impressive task breadth of recent models, critical gaps remain between what radiologists need and what current systems deliver:

  1. No calibrated confidence for diagnostic decisions. Generative VLMs predict diagnoses as free text, but they typically do not provide calibrated confidence scores. In clinical settings, confidence matters. Clinicians cannot tune sensitivity–specificity trade-offs across clinical contexts—an important requirement for real-world deployment. Discriminative models provide these properties but lack the flexibility of open-ended generation.
  2. Cross-entropy loss does not optimize clinical fidelity. Standard training methods treat all token-level errors similarly, regardless of their clinical consequences. A coordinate mistake may be penalized no more than a harmless wording change. A “yes” can be flipped to a “no” even though the clinical meaning is completely different. Missing a life-threatening finding may carry the same training penalty as omitting a minor observation. As a result, models are not explicitly optimized for what matters most in patient care.
  3. No capability for measurement-dependent findings. Some radiological findings require more than visual recognition. Radiological signs such as cardiomegaly, mediastinal widening etc. depend on precise measurements. For example, a model may correctly recognize whether a chest radiograph was acquired using an AP or PA view. But determining cardiomegaly requires measuring the cardiac and thoracic widths and determining the cardiothoracic ratio. Those quantities should be measured and computed rather than visually approximated while considering variables such as type of view, exposure, rotation of the patient etc. 

Together, these gaps call for more than a fluent generative model. The system must combine broad task coverage, structured predictions, clinically aligned optimization, and quantitative tools where direct measurement is required.

CARE-X: One model, flexible outputs

CARE-X brings these diverse interpretation capabilities into one model, using generative or dual inference according to the needs of each task:

Task typeWhat CARE-X doesInference modeReport generation: FindingsProduces the detailed findings sectionGenerativeReport generation: ImpressionProduces the concise diagnostic impressionGenerativePresence and negation assessmentDetermines whether a pathology is present or absent and handles negationDual: generative + auxiliary headDisease location assessmentIdentifies where an abnormality appearsGenerativeFine-grained multilabel disease classificationCategorizes abnormalities across multiple labelsGenerativeMultilabel tubes and lines classificationIdentifies visible medical devicesGenerativeAbnormal placement detection of tubes and linesDetermines whether a device is positioned incorrectlyDual: generative + auxiliary headAbnormality phrase groundingLocalizes a described pathological findingDual: generative + auxiliary headAnatomical groundingLocalizes 29 anatomical regionsDual: generative + auxiliary headTable 1: CARE-X task coverage and inference modes

Dual inference means that a single forward pass produces both an autoregressive response and a structured auxiliary-head prediction with a confidence score. This provides free-text flexibility alongside threshold-adjustable outputs for tasks where operating-point control matters.

The CARE-X architecture and training approach

CARE-X is built on a SigLIP2-so400M vision encoder and a Phi-4-mini-instruct (3.8B) language model connected through a lightweight adapter. To support both free-text generation and structured clinical predictions, the model augments the shared language backbone with task-specific auxiliary heads for classification and visual grounding. These heads provide calibrated diagnostic predictions and spatial localization signals while sharing representations with the generative language model. Rather than being trained independently, they are co-trained with the language-modeling objective, allowing structured supervision to enrich shared representations and improve generative performance on the same tasks.

Training. CARE-X uses a three-stage supervised fine-tuning pipeline (vision pre-training, adapter/head training, and LoRA adaptation) followed by DAPO-based reinforcement learning. DAPO optimizes task-specific rewards for clinical reporting, diagnostic accuracy, and spatial grounding quality.

Figure 1. The CARE-X model. (Left) Supervised fine-tuning with task-specific heads — classification, grounding, and language modeling — sharing the same Phi-4-mini-instruct backbone. The classification head outputs calibrated P(Yes)/P(No) scores; the grounding head outputs bounding box coordinate with confidence; the language modeling head generates free-text responses. (Right) DAPO with task-specific rewards for multi-task reinforcement alignment across report generation, grounding, and VQA.  Auxiliary supervision: Structured prediction strengthens generation

A central finding of this work is that co-training discriminative auxiliary heads with a generative VLM enriches shared representations, leading to stronger generative performance on the same tasks while also providing calibrated structured predictions.

Grounding improvements

The auxiliary grounding head consistently improves localization over generative decoding. On anatomical grounding (Chest ImaGenome), mAP and mIoU increase by +28.2 pp and +6.2 pp, while the largest gains occur on phrase grounding (PadChest), with +24.6 pp mAP and +14.1 pp mIoU. The composite spatial loss enhances geometric precision in shared representations.

DAPO bridges the gap to dedicated detection heads

DAPO-trained generative output approaches or exceeds the SFT auxiliary detection head. On Anatomy grounding, CARE-X generative (0.868 mAP) surpasses the SFT detection head (0.865). This is practically significant—it demonstrates that reward-aligned learning can bring autoregressive spatial decoding to parity with structured prediction, offering clinicians a single generative inference mode without requiring auxiliary heads at test time.

Calibrated classification with tunable operating points

Beyond representation enrichment, the classification head offers a distinct deployment advantage: calibrated probability scores with tunable thresholds allow clinicians to shift between high-sensitivity screening and high-specificity confirmation from a single forward pass—a capability purely generative architectures cannot provide.

ModelInference SettingSensitivity ↑PPV ↑F1 ↑CARE-XGenerative0.9320.8950.913CARE-X (Th=0.5)Auxiliary Head0.9430.8850.913CARE-X (Th=0.6)Auxiliary Head0.8550.9270.890CheXOneGenerative0.8780.8540.866MedGemmaGenerative0.7980.8860.839Table 2: Abnormality classification performance on Chest ImaGenome. Adjustable thresholds enable operating-point selection. Strong report generation across four benchmarks

Within the paper’s comparison set, CARE-X achieves the strongest performance on most reported metrics across MIMIC-CXR, IU-Xray, CheXpert-Plus, and ReXGradient. CRIMSON, a held-out metric that evaluates abnormal findings and weights errors by clinical severity, suggests these gains reflect clinically meaningful improvements rather than reward-specific optimization.

Figure 2. CRIMSON scores (↑) for CARE-X against baseline report-generation models across four chest X-ray datasets — ReXGradient, MIMIC-CXR, IU-Xray, and CheXpert-Plus. CARE-X (highlighted) achieves the highest CRIMSON score on every dataset. CARE-X reaches 94% accuracy on ReXVQA

CARE-X ranks first on the ReXrank RexVQA leaderboard (opens in new tab) as of August 2026. On the ReXVQA benchmark (41,007 question–answer pairs across five clinically relevant categories), CARE-X reaches 94% overall accuracy, six percentage points above the next-best publicly reported model. 

Figure 3: ReXVQA accuracy across five findings-quality dimensions — negation, presence, location, differential diagnosis, geometric information, and overall. CARE-X consistently outperforms CheXOne-R1 and MedGemma on every axis, with the largest margins in differential diagnosis, location assessment and negation. Tool-augmented measurement: Interleaving perception and computation

Some radiological findings depend on quantitative measurements rather than visual patterns. In a separate research experiment from CARE-X, we built an inference-time pipeline that combines Qwen3-VL-4B-Instruct with deterministic measurement tools, allowing the model to alternate between image understanding and precise computation. Qwen3-VL-4B-Instruct retains visual access to the radiograph throughout inference, invoking tools to identify anatomical landmarks, compute measurements, and evaluate diagnostic thresholds as needed. This creates a multi-turn reasoning loop that interleaves perception and measurement, enabling the model to combine visual context with exact quantitative evidence before reaching a diagnosis.

Figure 4. Tool-augmented quantitative reasoning pipeline. The orchestrator mediates a multi-turn loop: the VLM reasons over the image (perception), emits structured tool calls, receives deterministic results, and synthesizes the final diagnosis.

Despite requiring no task-specific training, this approach substantially outperforms perception-only inference across all evaluated measurement-based conditions. The results suggest that for threshold-dependent diagnoses, direct computation of clinically defined measurements is more reliable than visual approximation alone.

More broadly, this measurement-augmented approach could augment clinical workflows by expanding the set of quantitative assessments routinely derived from chest radiographs. For example, aortic dilation is not typically quantified on CXR and is often detected only incidentally on CT scans obtained for other indications. As delayed detection can contribute to adverse cardiovascular outcomes, reliable CXR-based screening could enable earlier identification and follow-up of aortic dilation.

ConditionPerception F1Tool F1Δ F1Cardiomegaly74.5696.00+21.4Mediastinal Widening72.6397.47+24.8Aortic Knob Enlargement60.3199.76+39.5Ascending Aorta Enlargement39.33100.00+60.7Descending Aorta Enlargement†28.57100.00+71.4Average+43.6Table 3: Perception-only versus tool-augmented measurement. The average F1 improvement is 43.6 percentage points across five conditions. Validation on Indian clinical data: Rare ICU conditions and CT-confirmed enlargement

Research ethics and data use: The Narayana Health evaluations used de-identified, retrospective clinical data under applicable institutional ethics review and data-use approvals. Narayana Health approved publication of the study results described here. 

Study 1: Inpatient and ICU conditions

To assess real-world generalizability in a research setting, we evaluated CARE-X on 1,047 de-identified chest radiographs from Narayana Health, annotated for five rare, high-acuity conditions with prevalence ranging from 2.6% to 5.2%—reflecting realistic clinical distributions where missed diagnoses carry severe consequences. 

FractureMediastinal ShiftPneumoperitoneumPneumothoraxTubes & Lines Abnormal PlacementModelSens / SpecSens / SpecSens / SpecSens / SpecSens / SpecCheXOne0.41 / 0.900.80 / 0.780.67 / 0.980.85 / 0.720.03 / 0.97MedGemma0.05 / 1.001.00 / 0.530.00 / 1.000.52 / 0.730.18 / 0.87CARE-X0.62 / 0.640.83 / 0.860.89 / 0.940.83 / 0.750.66 / 0.77Table 4: ICU pathology classification on Indian hospital data. CARE-X achieves the most balanced performance.

CARE-X achieves the highest sensitivity in three out of five conditions while maintaining reasonable specificity, demonstrating generalization to low-prevalence clinical settings.

Study 2: CT-confirmed enlargement conditions

In a retrospective study to measure pure recall efficacy, we evaluated measurement-dependent conditions such as mediastinal widening findings including aortic enlargement, hilar mass, and pulmonary artery enlargement on a outpatient cohort of 122 positive cases with CT-confirmed ground truth, avoiding the subjectivity of radiologist consensus on borderline enlargement findings on CXR. In the overlay setting, the VLM receives the original radiograph alongside a second image with condition-relevant anatomical segmentation masks — offering spatial guidance without direct access to measurement tools.

The tool-augmented variant reached 94.26% recall, a +10.65 percentage-point gain over the best perception-only baseline. Where CT or echocardiography access is limited, reliable triage from a widely available modality like chest X-ray can cut both unnecessary referrals and missed diagnoses.

Figure 5: Recall on the CT-confirmed enlargement cohort across perception-only, overlay-assisted, and tool-augmented inference. (Study 2)

In a related study (accepted at EACTS conference 2026), for mild aortic dilation, the measurement-driven reasoning approach detected 40 of 43 CT-confirmed cases (93% sensitivity), compared to just 5 of 43 (12%) identified on the initial radiology reads, where aortic enlargement is usually not the primary indication for the chest X-ray. This corresponds to 35 additional mild cases that were surfaced but missed during the initial CXR interpretation. These results suggest that explicit quantitative measurements may help identify borderline enlargement that is difficult to assess through visual inspection alone. 

What this does and doesn’t show 

These numbers are all recall, i.e., how many true positives we catch. This was the focus of the initial study because, in triage, a missed diagnosis is typically the costlier failure mode, and CT-confirmed ground truth gave us a clean way to measure it without relying on radiologist consensus for the difficult cases. 

Recall, however, captures only one dimension of diagnostic performance. A model that flags everything achieves perfect recall and is useless in practice. An extended study is underway that includes CT-confirmed negative cohorts as well. Preliminary results are promising, and further studies are planned to explicitly evaluate the viability of quantitative aortic measurements on chest X-ray as a screening tool for aortic dilation. 

CARE-X: Toward clinically useful radiology AI

CARE-X demonstrates that discriminative and generative objectives can be effectively combined within a unified radiology AI model. By jointly training classification, grounding, and language capabilities, the model supports both flexible report generation and calibrated, threshold-adjustable predictions. The separate measurement study further highlights a practical division of labor between learned reasoning and deterministic computation: the VLM provides visual understanding and identifies relevant evidence, while measurement-dependent diagnoses are computed through transparent, tool-based calculations. Retrospective evaluation on clinically challenging Narayana Health cohorts provides encouraging evidence of the potential of this approach for real-world radiology applications. The clinical relevance of this research is underscored by the selection of the AI-based aortic dilatation screening application as a finalist for showcase at the IHF Innovation Hub, World Hospital Congress 2026, recognizing its potential to support earlier detection and clinical decision-making in cardiovascular care. 

Looking ahead, CARE-X can be extended beyond its current capabilities through structured report generation, richer differential diagnosis support, and tighter integration of tools within the model itself. The framework could also benefit from incorporating broader clinical context, including laboratory results and patient history, enabling more comprehensive clinical reasoning. 

CARE-X is a research model, not a Microsoft product offering or medical device. It has not been cleared or approved by any regulatory authority and is not intended or validated for clinical diagnosis, screening, patient care, or clinical decision-making. The results described are retrospective research findings and do not establish safety, effectiveness, or suitability for clinical use.  

Paper co-authors:

Mercy Ranjit, Anirban Porya (opens in new tab), Niharika Vadlamudi (opens in new tab), Nikhilesh E (opens in new tab), Sathvik Joel (opens in new tab), Prasanth V V (opens in new tab), Tanuja Ganu, Abhyuday Swamy (opens in new tab), Pranay Umredkar (opens in new tab), Pradeep Narayan (opens in new tab), Vivek Rajagopal (opens in new tab)

Collaborators: Medha AI (opens in new tab), Narayana Health (opens in new tab)

Opens in a new tab

The post Introducing CARE-X: Towards Clinically Useful Radiology VLMs with Auxiliary Supervision, Reward-Aligned Learning, and Tool-Augmented Measurement appeared first on Microsoft Research.

Categories: Microsoft

Introducing CARE-X: Towards Clinically Useful Radiology VLMs with Auxiliary Supervision, Reward-Aligned Learning, and Tool-Augmented Measurement

Microsoft Research - Tue, 08/11/2026 - 18:00

Research Note: CARE-X is a research model and not a Microsoft product offering or medical device. It has not been cleared or approved by any regulatory authority and is not intended for clinical diagnosis, screening, or patient care. The results described below are retrospective research findings and do not establish the safety, effectiveness, or suitability of CARE-X for any clinical use. References to potential workflows describe areas for future research, not currently available capabilities or recommended uses. 

At a glance
  • The challenge: Chest X-ray interpretation spans diverse tasks that require both expressive report generation and calibrated diagnostic predictions.
  • CARE-X is a unified chest X-ray VLM for diverse clinical interpretation tasks. It combines generation and structured prediction to provide both free-text reasoning and deterministic outputs.
  • CARE-X uses reinforcement learning (DAPO) to reward clinical correctness in a multi-task setting.
  • In a separate research experiment from CARE-X, we paired Qwen3-VL-4B-Instruct with deterministic measurement tools to evaluate whether direct computation could improve performance on measurement-dependent conditions compared with visual approximation alone. 
  • Validated on real-world Indian clinical data from Narayana Health, including rare ICU pathologies and CT-confirmed enlargement conditions.
What radiologists need: Task diversity, flexibility, and clinical fidelity

A clinically useful radiology AI system must support a wide range of tasks, adapt to different workflows, and produce outputs that are medically accurate.

Radiologists and other clinicians use chest X-rays for many different purposes. A clinically useful AI system must be able to support that range of tasks. It may be asked to generate detailed findings and concise impressions for a report, answer questions about the presence, absence, or location of a finding, identify medical devices and assess their placement, or pinpoint exactly where an abnormality appears in an image.

These tasks also require different kinds of outputs, from narrative reports to calibrated diagnostic scores. And above all, they require clinical accuracy. A report could ostensibly be perfectly written yet clinically wrong if it misses a finding, reverses a negation, or misidentifies a location. Certain findings could be trivial in one context and vital to identify in another.

CARE-X was developed as a research model to explore how a unified approach can address these diverse demands. The system combines generative and discriminative capabilities, clinically aligned optimization, and tool-based reasoning to support a broader range of radiology workflows while maintaining clinical fidelity.

Azure AI Foundry Labs

Get a glimpse of potential future directions for AI, with these experimental technologies from Microsoft Research.

Azure AI Foundry Opens in a new tab Gaps in current radiology vision-language models

Despite the impressive task breadth of recent models, critical gaps remain between what radiologists need and what current systems deliver:

  1. No calibrated confidence for diagnostic decisions. Generative VLMs predict diagnoses as free text, but they typically do not provide calibrated confidence scores. In clinical settings, confidence matters. Clinicians cannot tune sensitivity–specificity trade-offs across clinical contexts—an important requirement for real-world deployment. Discriminative models provide these properties but lack the flexibility of open-ended generation.
  2. Cross-entropy loss does not optimize clinical fidelity. Standard training methods treat all token-level errors similarly, regardless of their clinical consequences. A coordinate mistake may be penalized no more than a harmless wording change. A “yes” can be flipped to a “no” even though the clinical meaning is completely different. Missing a life-threatening finding may carry the same training penalty as omitting a minor observation. As a result, models are not explicitly optimized for what matters most in patient care.
  3. No capability for measurement-dependent findings. Some radiological findings require more than visual recognition. Radiological signs such as cardiomegaly, mediastinal widening etc. depend on precise measurements. For example, a model may correctly recognize whether a chest radiograph was acquired using an AP or PA view. But determining cardiomegaly requires measuring the cardiac and thoracic widths and determining the cardiothoracic ratio. Those quantities should be measured and computed rather than visually approximated while considering variables such as type of view, exposure, rotation of the patient etc. 

Together, these gaps call for more than a fluent generative model. The system must combine broad task coverage, structured predictions, clinically aligned optimization, and quantitative tools where direct measurement is required.

CARE-X: One model, flexible outputs

CARE-X brings these diverse interpretation capabilities into one model, using generative or dual inference according to the needs of each task:

Task typeWhat CARE-X doesInference modeReport generation: FindingsProduces the detailed findings sectionGenerativeReport generation: ImpressionProduces the concise diagnostic impressionGenerativePresence and negation assessmentDetermines whether a pathology is present or absent and handles negationDual: generative + auxiliary headDisease location assessmentIdentifies where an abnormality appearsGenerativeFine-grained multilabel disease classificationCategorizes abnormalities across multiple labelsGenerativeMultilabel tubes and lines classificationIdentifies visible medical devicesGenerativeAbnormal placement detection of tubes and linesDetermines whether a device is positioned incorrectlyDual: generative + auxiliary headAbnormality phrase groundingLocalizes a described pathological findingDual: generative + auxiliary headAnatomical groundingLocalizes 29 anatomical regionsDual: generative + auxiliary headTable 1: CARE-X task coverage and inference modes

Dual inference means that a single forward pass produces both an autoregressive response and a structured auxiliary-head prediction with a confidence score. This provides free-text flexibility alongside threshold-adjustable outputs for tasks where operating-point control matters.

The CARE-X architecture and training approach

CARE-X is built on a SigLIP2-so400M vision encoder and a Phi-4-mini-instruct (3.8B) language model connected through a lightweight adapter. To support both free-text generation and structured clinical predictions, the model augments the shared language backbone with task-specific auxiliary heads for classification and visual grounding. These heads provide calibrated diagnostic predictions and spatial localization signals while sharing representations with the generative language model. Rather than being trained independently, they are co-trained with the language-modeling objective, allowing structured supervision to enrich shared representations and improve generative performance on the same tasks.

Training. CARE-X uses a three-stage supervised fine-tuning pipeline (vision pre-training, adapter/head training, and LoRA adaptation) followed by DAPO-based reinforcement learning. DAPO optimizes task-specific rewards for clinical reporting, diagnostic accuracy, and spatial grounding quality.

Figure 1. The CARE-X model. (Left) Supervised fine-tuning with task-specific heads — classification, grounding, and language modeling — sharing the same Phi-4-mini-instruct backbone. The classification head outputs calibrated P(Yes)/P(No) scores; the grounding head outputs bounding box coordinate with confidence; the language modeling head generates free-text responses. (Right) DAPO with task-specific rewards for multi-task reinforcement alignment across report generation, grounding, and VQA.  Auxiliary supervision: Structured prediction strengthens generation

A central finding of this work is that co-training discriminative auxiliary heads with a generative VLM enriches shared representations, leading to stronger generative performance on the same tasks while also providing calibrated structured predictions.

Grounding improvements

The auxiliary grounding head consistently improves localization over generative decoding. On anatomical grounding (Chest ImaGenome), mAP and mIoU increase by +28.2 pp and +6.2 pp, while the largest gains occur on phrase grounding (PadChest), with +24.6 pp mAP and +14.1 pp mIoU. The composite spatial loss enhances geometric precision in shared representations.

DAPO bridges the gap to dedicated detection heads

DAPO-trained generative output approaches or exceeds the SFT auxiliary detection head. On Anatomy grounding, CARE-X generative (0.868 mAP) surpasses the SFT detection head (0.865). This is practically significant—it demonstrates that reward-aligned learning can bring autoregressive spatial decoding to parity with structured prediction, offering clinicians a single generative inference mode without requiring auxiliary heads at test time.

Calibrated classification with tunable operating points

Beyond representation enrichment, the classification head offers a distinct deployment advantage: calibrated probability scores with tunable thresholds allow clinicians to shift between high-sensitivity screening and high-specificity confirmation from a single forward pass—a capability purely generative architectures cannot provide.

ModelInference SettingSensitivity ↑PPV ↑F1 ↑CARE-XGenerative0.9320.8950.913CARE-X (Th=0.5)Auxiliary Head0.9430.8850.913CARE-X (Th=0.6)Auxiliary Head0.8550.9270.890CheXOneGenerative0.8780.8540.866MedGemmaGenerative0.7980.8860.839Table 2: Abnormality classification performance on Chest ImaGenome. Adjustable thresholds enable operating-point selection. Strong report generation across four benchmarks

Within the paper’s comparison set, CARE-X achieves the strongest performance on most reported metrics across MIMIC-CXR, IU-Xray, CheXpert-Plus, and ReXGradient. CRIMSON, a held-out metric that evaluates abnormal findings and weights errors by clinical severity, suggests these gains reflect clinically meaningful improvements rather than reward-specific optimization.

Figure 2. CRIMSON scores (↑) for CARE-X against baseline report-generation models across four chest X-ray datasets — ReXGradient, MIMIC-CXR, IU-Xray, and CheXpert-Plus. CARE-X (highlighted) achieves the highest CRIMSON score on every dataset. CARE-X reaches 94% accuracy on ReXVQA

CARE-X ranks first on the ReXrank RexVQAleaderboard (opens in new tab) as of August 2026. On the ReXVQA benchmark (41,007 question–answer pairs across five clinically relevant categories), CARE-X reaches 94% overall accuracy, six percentage points above the next-best publicly reported model. 

Figure 3: ReXVQA accuracy across five findings-quality dimensions — negation, presence, location, differential diagnosis, geometric information, and overall. CARE-X consistently outperforms CheXOne-R1 and MedGemma on every axis, with the largest margins in differential diagnosis, location assessment and negation. Tool-augmented measurement: Interleaving perception and computation

Some radiological findings depend on quantitative measurements rather than visual patterns. In a separate research experiment from CARE-X, we built an inference-time pipeline that combines Qwen3-VL-4B-Instruct with deterministic measurement tools, allowing the model to alternate between image understanding and precise computation. Qwen3-VL-4B-Instruct retains visual access to the radiograph throughout inference, invoking tools to identify anatomical landmarks, compute measurements, and evaluate diagnostic thresholds as needed. This creates a multi-turn reasoning loop that interleaves perception and measurement, enabling the model to combine visual context with exact quantitative evidence before reaching a diagnosis.

Figure 4. Tool-augmented quantitative reasoning pipeline. The orchestrator mediates a multi-turn loop: the VLM reasons over the image (perception), emits structured tool calls, receives deterministic results, and synthesizes the final diagnosis.

Despite requiring no task-specific training, this approach substantially outperforms perception-only inference across all evaluated measurement-based conditions. The results suggest that for threshold-dependent diagnoses, direct computation of clinically defined measurements is more reliable than visual approximation alone.

More broadly, this measurement-augmented approach could augment clinical workflows by expanding the set of quantitative assessments routinely derived from chest radiographs. For example, aortic dilation is not typically quantified on CXR and is often detected only incidentally on CT scans obtained for other indications. As delayed detection can contribute to adverse cardiovascular outcomes, reliable CXR-based screening could enable earlier identification and follow-up of aortic dilation.

ConditionPerception F1Tool F1Δ F1Cardiomegaly74.5696.00+21.4Mediastinal Widening72.6397.47+24.8Aortic Knob Enlargement60.3199.76+39.5Ascending Aorta Enlargement39.33100.00+60.7Descending Aorta Enlargement†28.57100.00+71.4Average+43.6Table 3: Perception-only versus tool-augmented measurement. The average F1 improvement is 43.6 percentage points across five conditions. Validation on Indian clinical data: Rare ICU conditions and CT-confirmed enlargement

Research ethics and data use: The Narayana Health evaluations used de-identified, retrospective clinical data under applicable institutional ethics review and data-use approvals. Narayana Health approved publication of the study results described here. 

Study 1: Inpatient and ICU conditions

To assess real-world generalizability in a research setting, we evaluated CARE-X on 1,047 de-identified chest radiographs from Narayana Health, annotated for five rare, high-acuity conditions with prevalence ranging from 2.6% to 5.2%—reflecting realistic clinical distributions where missed diagnoses carry severe consequences. 

FractureMediastinal ShiftPneumoperitoneumPneumothoraxTubes & Lines Abnormal PlacementModelSens / SpecSens / SpecSens / SpecSens / SpecSens / SpecCheXOne0.41 / 0.900.80 / 0.780.67 / 0.980.85 / 0.720.03 / 0.97MedGemma0.05 / 1.001.00 / 0.530.00 / 1.000.52 / 0.730.18 / 0.87CARE-X0.62 / 0.640.83 / 0.860.89 / 0.940.83 / 0.750.66 / 0.77Table 4: ICU pathology classification on Indian hospital data. CARE-X achieves the most balanced performance.

CARE-X achieves the highest sensitivity in three out of five conditions while maintaining reasonable specificity, demonstrating generalization to low-prevalence clinical settings.

Study 2: CT-confirmed enlargement conditions

In a retrospective study to measure pure recall efficacy, we evaluated measurement-dependent conditions such as mediastinal widening findings including aortic enlargement, hilar mass, and pulmonary artery enlargement on a outpatient cohort of 122 positive cases with CT-confirmed ground truth, avoiding the subjectivity of radiologist consensus on borderline enlargement findings on CXR. In the overlay setting, the VLM receives the original radiograph alongside a second image with condition-relevant anatomical segmentation masks — offering spatial guidance without direct access to measurement tools.

The tool-augmented variant reached 94.26% recall, a +10.65 percentage-point gain over the best perception-only baseline. Where CT or echocardiography access is limited, reliable triage from a widely available modality like chest X-ray can cut both unnecessary referrals and missed diagnoses.

Figure 5: Recall on the CT-confirmed enlargement cohort across perception-only, overlay-assisted, and tool-augmented inference. (Study 2)

In a related study (accepted at EACTS conference 2026), for mild aortic dilation, the measurement-driven reasoning approach detected 40 of 43 CT-confirmed cases (93% sensitivity), compared to just 5 of 43 (12%) identified on the initial radiology reads, where aortic enlargement is usually not the primary indication for the chest X-ray. This corresponds to 35 additional mild cases that were surfaced but missed during the initial CXR interpretation. These results suggest that explicit quantitative measurements may help identify borderline enlargement that is difficult to assess through visual inspection alone. 

What this does and doesn’t show 

These numbers are all recall, i.e., how many true positives we catch. This was the focus of the initial study because, in triage, a missed diagnosis is typically the costlier failure mode, and CT-confirmed ground truth gave us a clean way to measure it without relying on radiologist consensus for the difficult cases. 

Recall, however, captures only one dimension of diagnostic performance. A model that flags everything achieves perfect recall and is useless in practice. An extended study is underway that includes CT-confirmed negative cohorts as well. Preliminary results are promising, and further studies are planned to explicitly evaluate the viability of quantitative aortic measurements on chest X-ray as a screening tool for aortic dilation. 

CARE-X: Toward clinically useful radiology AI

CARE-X demonstrates that discriminative and generative objectives can be effectively combined within a unified radiology AI model. By jointly training classification, grounding, and language capabilities, the model supports both flexible report generation and calibrated, threshold-adjustable predictions. The separate measurement study further highlights a practical division of labor between learned reasoning and deterministic computation: the VLM provides visual understanding and identifies relevant evidence, while measurement-dependent diagnoses are computed through transparent, tool-based calculations. Retrospective evaluation on clinically challenging Narayana Health cohorts provides encouraging evidence of the potential of this approach for real-world radiology applications. The clinical relevance of this research is underscored by the selection of the AI-based aortic dilatation screening application as a finalist for showcase at the IHF Innovation Hub, World Hospital Congress 2026, recognizing its potential to support earlier detection and clinical decision-making in cardiovascular care. 

Looking ahead, CARE-X can be extended beyond its current capabilities through structured report generation, richer differential diagnosis support, and tighter integration of tools within the model itself. The framework could also benefit from incorporating broader clinical context, including laboratory results and patient history, enabling more comprehensive clinical reasoning. 

CARE-X is a research model, not a Microsoft product offering or medical device. It has not been cleared or approved by any regulatory authority and is not intended or validated for clinical diagnosis, screening, patient care, or clinical decision-making. The results described are retrospective research findings and do not establish safety, effectiveness, or suitability for clinical use.  

Paper co-authors:

Mercy Ranjit, Anirban Porya (opens in new tab), Niharika Vadlamudi (opens in new tab), Nikhilesh E (opens in new tab), Sathvik Joel (opens in new tab), Prasanth V V (opens in new tab), Tanuja Ganu, Abhyuday Swamy (opens in new tab), Pranay Umredkar (opens in new tab), Pradeep Narayan (opens in new tab), Vivek Rajagopal (opens in new tab)

Collaborators: Medhai AI, Narayana Health (opens in new tab)

Opens in a new tab

The post Introducing CARE-X: Towards Clinically Useful Radiology VLMs with Auxiliary Supervision, Reward-Aligned Learning, and Tool-Augmented Measurement appeared first on Microsoft Research.

Categories: Microsoft

Orchard: An open framework for scalable agentic AI

Microsoft Research - Mon, 08/03/2026 - 18:00
At a glance
  • Orchard is an open-source framework for scalable and cost-effective agentic AI research, built around Orchard Env, a reusable environment service for training and evaluating agents across task domains.
  • The same Orchard infrastructure supports software-engineering, web-navigation, and personal-assistant agents, and can train them directly inside real deployment harnesses such as Codex, OpenClaw, and ZeroClaw—letting researchers reuse environments, data pipelines, and evaluation workflows across tasks. 
  • Orchard-SWE, Orchard-GUI, and Orchard-Claw demonstrate that relatively small open-weight models can achieve strong results on complex real-world tasks. For example, Orchard-SWE reaches 69.7% on SWE-bench Verified—73.0% with value-model reranking—using only about 3 billion active parameters, approaching frontier systems using more than 10 times larger models. 

Alongside the models and workflows, the project releases training data and evaluation methods intended to help the broader research community build and study open agentic systems. Artificial intelligence is rapidly moving beyond static question-answering toward autonomous agents that can plan, reason, and act across complex, multistep environments. These systems can fix bugs in complex codebases, navigate the web on a user’s behalf, and manage workflows involving calendars and email.  

While there is excitement around agentic AI’s capabilities, the research community faces a persistent bottleneck. Building state-of-the-art agentic systems often requires proprietary infrastructure, including custom sandboxes, closed training pipelines, and proprietary datasets that most researchers and practitioners cannot access or reproduce.

To address this gap, we introduce Orchard (opens in new tab), an open-source framework for scalable agentic modeling. At the center of Orchard is Orchard Env, a lightweight, Kubernetes environment that provides reusable isolated components for running and building agents at scale—from collecting training data to reinforcement learning rollouts and evaluation.  

Unlike many existing frameworks, Orchard Env is designed to support different agent systems and task types without modification. The same service can support software-engineering agents, web-browsing agents, and personal-assistant agents across domains. 

To demonstrate this approach, we are releasing three domain-specific training recipes—Orchard-SWE, Orchard-GUI, and Orchard-Claw. (opens in new tab) We are also releasing the training data and evaluation methods used to build them.

Foundry Labs

Foundry Labs bridges the distance between Microsoft breakthroughs and the builders, scientists, and curious minds who can put them to work.

Foundry Labs Opens in a new tab Environment layer that scales across types of tasks

The central idea behind Orchard is that the runtime environment should be a standalone, reusable service rather than infrastructure embedded inside a specific training framework. Orchard Env’s Kubernetes foundation enables it to create, manage, and remove thousands of isolated components in parallel.

The system is designed to work across tasks like coding, web browsing, using tools. It is also designed to work across different agent systems, along with stages of the training and evaluation process, including data distillation and reinforcement learning rollouts.

This flexibility makes Orchard practical at a research scale. Teams can introduce new benchmarks, agent systems, or training algorithms without rebuilding the underlying infrastructure from scratch.

Orchard also makes it possible to train agents inside any harness. Today’s most capable agents rarely run as a bare model. They operate through sophisticated harnesses—such as Claude Code, Codex, and OpenClaw—that manage multi-turn reasoning, tool use, and connections to external systems. Open training tools usually cannot handle these stateful, multi-process harnesses, forcing researchers to train on a simplified stand-in and then deploy in the real setting, which creates a mismatch. Orchard closes this gap: a lightweight proxy records the harness’s own model calls as training data while each rollout runs in its own container, so an agent can be trained end-to-end directly in the harness that it will be deployed with—OpenClaw, Codex, ZeroClaw, or others—and across several harnesses.

Orchard-SWE: Advancing open-source software engineering agents

Software engineering is one of the most demanding settings for autonomous agents. It requires multi-step reasoning over real codebases, tool use, and the ability to recover from mistakes. Orchard-SWE is our training workflow for this domain. It is built using the Mini-SWE-Agent framework, designed to autonomously solve software engineering tasks, and evaluated on the widely used SWE-bench Verified benchmark, which tests a model’s ability to navigate, diagnose, and repair real-world codebases.

To train the system, we distilled 107,000 agent interactions from two advanced open-weight models (MiniMax-M2.5 and Qwen3.5-397B) covering a broad range of GitHub Issues. The training process uses credit-assignment supervised fine-tuning: rather than discarding attempts where the agent failed to fully resolve an issue, the system learns from the productive portions of those partial attempts, expanding the amount of useful training data available to the model.

Reinforcement learning comes next, but its feedback is sparse—an agent usually learns only whether its final patch passed or failed the hidden tests. We start with Balanced Adaptive Rollout, designed to make the most of these infrequent success signals, and then add two “dense reward” techniques for richer guidance: on-policy distillation, in which a stronger teacher model scores the agent’s decisions step by step, and a process reward model, in which an AI judge rewards sound problem-solving process—writing tests that reproduce the bug, verifying the fix, and checking that existing behavior still works—independent of whether the final tests passed. 

Finally, we train a value model on past rollouts to rerank candidate solutions. Reinforcement learning generates many practice trajectories that are normally discarded; instead, trajectories from 20 prior experiments train a compact 4-billion-parameter value model that recognizes high-quality solutions, and at problem-solving time it scores several candidate answers and picks the best one. Together, these techniques take Orchard-SWE from a 61.4% baseline on SWE-bench Verified to 69.1% with Balanced Adaptive Rollout and 69.7% with the dense-reward techniques—a new state of the art among open-source models of comparable size (roughly 3 billion active parameters)—rising to 73% with value-model reranking, approaching frontier systems more than 10 times larger, as shown in Figure 1. 

Orchard-GUI: A lightweight browser agent for real-world web tasks

Web navigation presents a different set of challenges. Agents must interpret visual layouts, interact with dynamic interfaces, and complete open-ended tasks described only in natural language.

Orchard-GUI trains a 4-billion-parameter vision-language model as a browser agent using a relatively small amount of supervision: 400 distilled demonstrations combined with 2,200 open-ended training tasks. Despite this limited training data, the resulting model achieves strong results across several web-navigation benchmarks: 74.1% on WebVoyager, 67.0% on Online-Mind2Web, and 64.0% on DeepShop, for an average of 68.4%, as shown in Figure 1.

Figure 1. Performance comparison. Left: Orchard-SWE (35B-A3B, ~3B active) reaches 69.7% on SWE-bench Verified—73% with value-model reranking—matching frontier systems more than 10x larger. Right: Orchard-GUI (4B) achieves 68.4% average success across WebVoyager, Online-Mind2web, and DeepShop, making it the strongest open-source GUI agent while staying on par with proprietary systems from OpenAI and Google.

These results place Orchard-GUI among the strongest open-source web agents to date while remaining competitive with larger proprietary models. The results also suggest that with the right training approach and environment, small open models can perform well on real-world web tasks.

Orchard-Claw: Personal assistant agents for everyday productivity

Many of the most impactful agentic applications involve everyday productivity tasks, including reading and drafting emails, managing calendars, searching for information, and coordinating across tools. Orchard-Claw focuses on personal-assistant tasks by training an agent on just 200 synthetic tasks. Evaluated on Claw-Eval, a benchmark covering realistic productivity workflows, it successfully completes 59.6% of tasks when given up to three attempts. That increases to 73.9% when paired with the stronger ZeroClaw agent system.

Because Orchard can train agents directly inside real deployment harnesses, Orchard-Claw is trained across several of them—including ReACT, ZeroClaw, OpenClaw, and Codex—rather than a single simplified loop. Training inside these real harnesses substantially improves the agent’s reliability; under the Codex harness, for example, its success rate rises from 18.6% for the untrained model to 51.5% after Orchard training. 

Figure 2. Overview of the Orchard framework. Orchard Env (center) is a lightweight, Kubernetes-native environment service that provides shared capabilities such as sandbox management, command execution, file access, network controls, a REST API, and agent integration. It supports a range of task environments (bottom row) and is used across three task domains (top row): Orchard-SWE (software engineering), Orchard-GUI (browser navigation), and Orchard-Claw (AI personal assistant). Implications and the road ahead

Orchard’s results reinforce a broader point: the environment layer matters. By making the underlying infrastructure open, lightweight, and reusable, Orchard lowers the cost of agentic AI research. Teams no longer need to build custom isolated environments from scratch or depend on proprietary cloud services. The same Orchard Env can be used to generate training data, run reinforcement learning rollouts, and evaluate final models without rebuilding the system each time.

Looking ahead, we see reusing training experience as a promising direction toward cumulative agent learning. Instead of discarding trajectories once a training run finishes, we treat them as persistent assets—for example, distilling them into reusable value models. This enables agentic experience to accumulate over time, allowing each new generation of agents to inherit and extend the knowledge acquired by previous ones, rather than starting from scratch. 

The data efficiency demonstrated by Orchard-GUI suggests that larger-scale web agents could be trained without requiring large amounts of manually created training data. By releasing the complete Orchard stack, including the environment service, training pipelines, and training datasets, we hope to help the broader research community build more capable open agents more quickly. 

Acknowledgements

We thank the teams at Microsoft Research and collaborating institutions for their contributions to Orchard, as well as the open-source community whose benchmarks and tools made this research possible.

Opens in a new tab

The post Orchard: An open framework for scalable agentic AI appeared first on Microsoft Research.

Categories: Microsoft

Echoverse: Deep, evolving environments for computer-use agents

Microsoft Research - Thu, 07/30/2026 - 19:00
Scaling fidelity over sheer count, targeting the capabilities agents actually lack, and evolving with the models they train. At a glance

We built twelve training worlds for computer-use agents: ten deep domain worlds and two capability worlds, each drilling a single control rendered in many forms (date pickers and nested filters). Depth is what makes them worth training on: these worlds reproduce an application’s real behavior, come seeded with realistic data, and keep state coherent across screens and users. Trained on all twelve, a 9B model nearly doubles its base score (36.5% to 67.1%), coming within fourteen points of GPT-5.4. The experiment taught us several lessons: 

A computer-use agent learns the results of what its actions do only where they have real consequences. A click changes saved state, a message reaches a real person, or a page that refuses to move tells the agent its last move did nothing. A screenshot can show what an interface looks like, but only a working world shows what an action caused.

The consequences worth learning from are stateful, and most of them sit behind a login. The work people want automated lives in closed systems: email and chat, banking, health records, the internal consoles for cloud and ML. You cannot train an agent against the live versions of these. Every attempt writes to a real account, there is no reset between tries, and the true state stays hidden behind the screen. So you rebuild the system as a synthetic world where the database is yours: the state is real and changes for real, but it is safe to break, quick to reset, and graded from the data rather than a screenshot.

By a world we mean three things bound together: an environment (the application, its state, and the actions that change it), the tasks that set goals in it, and a verifier that grades the outcome against ground truth. The community is now good at making them: pipelines stand up an application, seed it, generate tasks, and attach verifiers, yielding hundreds of environments and thousands of checkable tasks. This work builds on that progress. However, once worlds are plentiful and its internal structure becomes the bottleneck: regardless of whether state stays coherent across users and screens, workflows keep their dependencies, a weak skill recurs in enough forms to generalize, and success is judged by outcome or by appearance.

Our bet, the one Echoverse tests, is that the real leverage comes less from adding worlds than from a loop that keeps improving the ones you already have. It treats building the environment and training the model as one process, not two stages: run a model in a world, find where it fails, make the world, its tasks, and its verifiers more faithful or more demanding there, train on the sharper signal, and repeat. Ordinary fine-tuning improves only the model. Here the same graded run that measures the model also improves the world that judged it, so a static benchmark saturates while the loop compounds.

Three levers keep that loop productive, none of them raw environment count. Depth: behaviorally faithful worlds for the domains that matter, including the closed and proprietary ones. Capability targeting: narrow worlds built around the exact interaction a model keeps failing. Co-evolution: improving the environment, its tasks, and its verifiers on every graded run, not just the model.

Figure 1: The learning loop: every graded run is read twice. Surviving failures become model training data, and defects in the world, its tasks, or its verifier become repairs. The same graded run that measures the model also sharpens the world. Why synthetic, and why deep?

Open, login-free sites might seem to remove the need for synthetic worlds, but they make a poor training ground for a different reason: they will not hold still. Pages get redesigned, listings and dates roll forward, and hosts throttle or block automated traffic, so a benchmark that is pinned to them drifts, and no two runs face the same site. An occasional evaluation can absorb that; training cannot, since it runs the same task thousands of times and needs the same world each time. A synthetic world is fixed in time and data: the calendar does not move, the seed data does not churn, and a task means the same thing on the thousandth rollout as on the first. We trade a little surface realism for a world we fully control.

Control is only the floor. A world can be perfectly stable and still be hollow, so what earns training time is depth: not its page count but how faithfully it preserves the causal structure of the work.   Five properties set the bar: behavioral fidelity (controls, permissions, and errors follow the product’s logic); coherent state (a sent message appears for its recipient, a cancelled meeting clears both calendars); workflow depth (an early choice constrains what happens later); authoritative verification (application state, not pixels); and domain value (the workflow is worth improving). In the systems that matter most, the difficulty lives in permissions, shared state, and audit histories: exactly the structure a shallow clone skips. Above this bar, more environments add variety; below it, they add noise.

How the Echoverse factory works?

Echoverse is a single pipeline with two outputs: full domain worlds that preserve workflow depth, and capability worlds that vary one diagnosed interaction. Both lean on the fact that we own the database underneath, so success is a property of the app’s own state, not a model’s read of a screenshot.

Building the world

The pipeline expands a handful of seed scenarios into a spec, then compiles it into machine-checkable claims about routes, state, and behavior. Only then does it generate the app: a FastAPI and SQLite backend under a React interface. A fresh app is a hypothesis, not a world: the builder runs every claim against the running environment, repairing the database, backend, or frontend until each passes, then writes a readiness record that separates hard blockers from advisory risks. A world with open blockers does not advance. Depth here is not a promise in a prompt; it is the list of claims the world has been shown to pass.

Growing the corpus

A world that builds cleanly is still not training data. We reground each task on the live database, drawing goals from entities that actually exist, then send every goal through a panel of analyzers: are its entities real, is the goal plausible, does its difficulty match the work, and, the sharpest test, can an agent driving the real UI complete it? That last check runs in the browser, catching goals no interface can satisfy before a model ever sees them. A generated goal is a claim; a solve against the real app is proof.

Every failure becomes an issue tagged by the layer that must change: database, backend, frontend, task text, or verifier. Layer-specific fixers apply the repair, re-check it against the running app, and roll it back if it regresses. The loop re-scores against database ground truth until the pass rate stops climbing, and each surviving task is exported carrying the exact check that grades it. Those tasks become training data through one process: GPT-5.4 solves each task, a verifier keeps the trajectories that pass ground truth, and those become the supervised fine-tuning (SFT) data behind every experiment below.

Building the world and growing the corpus are not two stages but rather one loop: most defects belong to the world, so we re-version the environment with every iteration. Harder tasks expose gaps in the world, and a sturdier world can carry harder tasks, so each round leaves both stronger.

Figure 2: The environment factory: the two loops behind every world. Phase 1 expands a handful of seeds into an app, then repairs the database, backend, and frontend until it passes machine-checkable claims. Phase 2 regrounds tasks on live data, runs a panel of analyzer and layer-specific fixer agents, and re-scores against database ground truth until the pass rate plateaus. Many of those fixes land in the world itself (dashed arrow).  The verifier is grounded in the database

Every task carries its own answer key, a value or a state change minted from the real database by a SQL query at generation, true by construction and re-checked after the agent finishes. A read is graded on semantic equivalence to the stored value ($288 for $287.62 passes); a write on a real before/after database diff, so claiming a ticket was closed fails unless the row flipped; a read_write scores the lower of the two. Grading is hard to game, grounded rather than labelled, and uniform across an EchoStay booking, an EchoForge issue, and an EchoBank transfer.

Full domains carry the workflow

The domains with the most consequential work are the hardest for public benchmarks to reach: closed, proprietary systems where the difficulty lives in permissions, shared state, and history, not layout. A faithful clone has to reproduce that. What matters is not the pixels but that an action’s consequences reach across screens and users, so a task can run a real workflow and be graded on the state it leaves behind.

The ten Echo domains span communication, technical work, regulated records, community, media, and travel. Where a rich public dataset exists we build on it: EchoStay is seeded from InsideAirbnb, so its listings, hosts, reviews, and amenities are real rather than invented, and EchoForum sits on a public forum corpus of 2.55 million comments. Where none exists, as with mail, calendar, banking, and health records, a seeding pipeline generates the state under strict constraints, dense and internally consistent, not a handful of placeholder rows.

Workflow categoryEnvironmentsDepth the world has to carryCommunication & coordinationEchoMail, EchoCalendar, EchoChatShared threads, schedules, participants, permissions, historiesTechnical creation & operationsEchoML, EchoForgeArtifacts, configuration, dependencies, roles, multi-stage changesRegulated records & transactionsEchoBank, EchoCareBalances or records, authorization, audit history, consequential writesCommunity, media & travelEchoForum, EchoTunes, EchoStayPersistent preferences, social state, search, booking, account actionsTable 1: The ten full-domain environments of the Echo family, grouped by the work they represent. Each is a faithful stand-in for a widely used product, named for the workflow rather than the brand.

That accumulated state is what makes an action’s consequences reach across screens and users. A booking in EchoStay moves through search, listing, availability, and payment across roughly 87 routes and 23 tables, but not a single confirmation screen; an EchoMail thread carries intent from draft through delivery, reply, and label state; an EchoCare order writes each change to an audit trail. The tasks are expensive because of it, often five to twenty actions deep, and finished only when the underlying state has changed.

Figure 3: Per-domain detail across the Echo suite. Each ships as a self-contained, fully-interactive clone of the app it models, with its own backend, seeded database, and feature surface. Counts are grounded database state, not mockups. Capability worlds isolate one skill

Not every weakness represents a missing domain; some are caused by a single control that the agent cannot reliably operate. Picture an agent booking a trip: it searches, filters, opens the right listing, then stalls at the date picker, unable to turn “the second week of March” into the right clicks on an unfamiliar calendar. Building another booking site would not fix that. The skill is learned only when the control itself appears in enough forms, and date pickers and nested filter-and-search are ubiquitous on the live web, rendered a hundred different ways, exactly the variability a single deep app cannot supply.

So we isolate the control and widen the interaction, mass-producing it across layouts, states, and constraints, then generating grounded tasks over each. The datepicker world renders one date control as six core widgets across 10 contexts and holds out 10 new unseen ones, from calendar heatmaps to scroll wheels and fiscal-quarter pickers; its hardest tasks turn transcription into reasoning, resolving “the last Thursday of January 2026” or “10 business days after a start date” to one exact, widget-reachable date. The nested-filter world varies 20 widget families and holds out nine compound-panel families as out-of-distribution, grading every submission by whether the filtered results actually meet the requested conditions, judged by the app’s own logic rather than by appearance.

Figure 4: Every widget family the two skills cover, split into training (in-distribution) and evaluation-only (held out): nested filters, 20 families plus 9 held-out compound panels; date pickers, 6 core types across 10 contexts plus 10 held-out widgets. Figure 5: Date pickers and nested filters themed across domains: nested filters over six verticals, from real estate to pet adoption; date pickers over ten contexts, from scheduling to insurance. What deeper, targeted worlds change

More trajectories do not automatically provide more training signal. What matters is depth: whether an episode carries a task through the dependent steps of a real workflow rather than just rehearsing an action in isolation. Two experiments make the difference concrete from opposite ends: one goes deeper on a whole domain, the other narrows to a single broken skill.

Shallow worlds backfire; deep worlds transfer

A shallow world is the cheap option. It stands up fast and looks convincing, but it only rehearses isolated, correct-looking clicks. Train on that and the model will pick up the wrong reflexes, over-stepping and looping and repeating dead actions, because nothing in the easy world ever punished them. A deep world costs more, but its trajectories carry the dependent structure that transfers to the live site.

To isolate that, take two live WebVoyager domains, Allrecipes and Hugging Face, and compare three checkpoints: the base model and two trained on shallow-world and deep-world trajectories built for those domains. The shallow world poses short, self-contained tasks; the deep world poses tasks that run across dependent steps, where an early action changes the state, options, and verification available later. Both give the model the same domain exposure, so only depth differs, and evaluation uses tasks from the public WebVoyager benchmark for these domains, run on the live sites outside any training world.

On Allrecipes, the shallow world pulls the model down, 80.0% to 75.0%; on Hugging Face it stays flat at 48.0%. Only the deep world improves both, lifting Allrecipes to 85.0% and the harder Hugging Face split to 65.0%. With exposure held equal, the gap is depth: the deep model loops less, and of the 37 Hugging Face tasks, those that exhaust their step budget fall from 15 to nine. What separated the two was not how much the model saw, but whether what it saw preserved the structure of the work.

Figure 6: Deep versus shallow worlds for two live WebVoyager domains, with identical domain exposure and different task depth. Deep lifts both; shallow drops below base on Allrecipes and stalls on Hugging Face. Precision about one skill

The datepicker and nested-filter worlds drill exactly the controls our evaluations flagged, and the two skills reinforce each other. Datepicker training lifts datepicker evaluations (in-distribution 60.0% to 82.6%, held-out layouts 34.0% to 54.0%); filter training lifts held-out filters 62.8% to 84.1%. Gains that hold on forms never trained on indicate that the model learned a rule, not a layout. The skills transfer across each other rather than competing: training either one alone still lifts the other, and training both is the best all-rounder on every split. Against GPT-5.4 as a frontier reference, that combined model already edges ahead on nested filters and closes most of the datepicker in-distribution gap, trailing clearly only on held-out datepickers. And the rule reaches the open web, lifting Online-Mind2Web 29.5% to 34.3% on sites it never saw. 

Figure 7: Targeted training, targeted gains: training either date pickers or nested filters lifts both controls, including held-out widgets and compositions neither was trained on, and training both is the best all-rounder on every split. Higher is better. From synthetic worlds to the live web

Three models run through the rest of this section. Base is Qwen3.5-9B given only a handful of synthetic trajectories, just enough to align a general model to the browser action space. Our model is that same 9-billion-parameter network trained on the full synthetic corpus. GPT-5.4 is a far larger frontier model, included as a reference ceiling.

Does the skill survive the open web? We evaluate our model, unchanged, on WebVoyager and Online-Mind2Web, benchmarks it never trained on. They barely overlap with what we built: both are dominated by open, public sites and read-mostly browsing, while our worlds train login-gated, write-heavy workflows. A large jump was never the point; direction is. The frozen model clears base on both, WebVoyager 66.5% to 71.5% and Online-Mind2Web 40.5% to 43.4% (without BrowserBase, 50.9% to 55.6% and 29.5% to 37.2%), reported through BrowserBase because a hosted browser strips the datacenter bot-blocks and rate limits that otherwise depress every agent’s score. With no live-web data in the mix, this is transfer, not memorization.

Figure 8: Synthetic training transfers to the live web. The full-corpus model, on two benchmarks it never trained on, clears base on both; scores run through BrowserBase to remove datacenter bot-blocks.

The modest live-web gain is a coverage effect, not a ceiling: aim at a live domain and it grows. EchoForge, our code-hosting world, is the same kind of app as GitHub, one of the live sites WebVoyager tests. Add EchoForge to the training mix and the live GitHub score climbs 58.5% to 63.4%, with the overall live scores rising too (WebVoyager 50.9% to 52.9%, Online-Mind2Web 29.5% to 31.1%). The average simply reflects that most of what we built sits in domains these benchmarks never touch.

Figure 9: Closing the gap to the frontier, per environment. The green bar is the gain from base to our model; the faded remainder is the distance still to GPT-5.4. Our model surpasses GPT-5.4 on EchoBank and both nested filters and closes most of the gap elsewhere; each model’s exact score is labelled on the right.

The domains we built, most of them closed and login-gated, tell the opposite story. Across all fourteen, the model nearly doubles base, 36.5% to 67.1%, and where base was weakest it climbs three- to nine-fold, with EchoCalendar, EchoML, EchoChat, EchoCare, EchoForge, and EchoForum all moving from single or low double digits into the forties through sixties. That puts a 9-billion-parameter model within fourteen points of GPT-5.4 on the average (67.1% against 80.7%). On EchoMail, EchoBank, and both nested filters, it matches or beats the far larger frontier model outright, trailing by only a few points on in-distribution datepickers. What gets a 9B model this close is not scale but training data that is deep, targeted, and checkable, exactly what the factory is built to produce. 

What scaling buys, and what it doesn’t

We scaled two axes separately: more trajectories through a fixed set of environments, drawn in equal numbers from each, and more distinct environments. They behave differently. More trajectories on the same worlds keep lifting the in-domain average, though the gains keep shrinking, while transfer to the live web flattens outright: from 6,400 to 20,000 trajectories, WebVoyager holds steady (54.8% to 55.6%) and Online-Mind2Web slips (40.1% to 37.2%). Since every point samples the worlds equally, this is no artifact: each environment holds only so much transferable skill, and once a model has drawn it out, more rollouts mostly polish what it already does. 

Scaling environments produces the opposite result. The average keeps climbing as breadth grows, and WebVoyager reaches its best only with the full set. For generalization, the lever is diversity, not volume. A model reaches sites it never saw by training across many kinds of work, not by seeing one kind many more times. 

Even so, scale itself is not the lever on either axis. A large trajectory budget spent on shallow worlds, or graded against the wrong answer, moves the synthetic number and goes nowhere on the live web. What travels is inside each trajectory: depth that preserves a real workflow, targeting that drills the control an agent fails, and database-grounded grading that keeps the signal honest.

Figure 10: Two scaling axes, scored without BrowserBase. Left: more trajectories on a fixed set of worlds, drawn in equal numbers from each, with the x-axis spaced by actual trajectory count. The synthetic average keeps rising, but live-web transfer saturates, WebVoyager flat and Online-Mind2Web slipping past 6,400 trajectories. Right: more environments, where breadth keeps the synthetic average and WebVoyager climbing. Diversity of environments, not trajectory volume, is what carries skill to unseen sites. The model is not the only thing that learns

The score an agent earns is never the model alone. It comes from a coupled stack: the agent, the environment, the task, and the verifier. A zero can mean the agent failed, or the control is broken, or the requested state is impossible, or the verifier checks the wrong thing. Reading every zero as model supervision trains on defects that should have been repaired. So, we read every graded rollout as a test of the whole stack and let the whole stack learn. The environment improves as broken controls and wiring get fixed, the tasks as goals are re-grounded and made harder, the verifier is fixed when it drifts out of sync with the data. Only failures that survive all three become model curriculum.

EchoStay made this visible. Its failures traced to the world, not the agent: a guest-count control silently broke booking tasks, so a correct booking could never register. Fixing it raised the share of those bookings that could be completed at all from 48% to 78%, recovering 15 of the 24 that had been blocked. The same loop finds different faults elsewhere: EchoForum needed frontend fixes and a page-load speedup, which took one failing set of 37 tasks from 0 solved to 36; EchoChat’s verifier had drifted out of sync with the data, and realigning it lifted the share of gradable tasks from 34% to 99%; EchoCare needed one state-wiring fix; EchoForge had the backend logic but no UI control to reach it.

As the world sharpens, the model climbs with it. Re-running the loop on EchoStay across two rounds, the model trained on its corpus more than doubles, from 16.2% to 38.5%, two-thirds of the distance to GPT-5.4’s 50.4%. The model is not the only thing that learns; it is the thing that compounds once everything under it learns.

Figure 11: Co-evolution lifts the model on EchoStay. As the world went from v1 to v2, the model trained on it more than doubled, from 16.2% to 38.5%, a separate measure from the world’s own solve rate. Higher is better. 

That boundary between repairing the world and teaching the model is easy to hold inside a controlled environment, where both are inspectable. The live web erases it: there is no world to repair mid-task, so when an action lands on nothing, correctness rests entirely on the agent noticing and choosing differently. That is the last thing a world has to teach, and where the live web is least forgiving.

From SFT to RL: Turning worlds into RLEs

Every result so far comes from imitation: the 9B model copies the trajectories GPT-5.4 got right. Imitation inherits a ceiling, though: a clean demonstration never shows how to recover from a mistake or when to stop, the failures that break agents in the wild. Reinforcement learning optimizes the outcome we grade and lets the model learn from its own trajectories, not a teacher’s.

But reinforcement learning needs an RL environment (RLE) it can drive at scale. Each rollout needs a reset to a known state, throughput to sample in parallel, and a reward it can trust, and a run replays the same task thousands of times. The live web is not an RLE: it will not reset, so no two rollouts begin alike; it throttles and blocks automated traffic well before RL’s scale; and it exposes no ground truth, only a screenshot a second model must judge, so the reward is as noisy as the judge and a policy learns the judge’s blind spots rather than the task. Echoverse is an RLE by construction. Every world is a self-contained app we snapshot and reset per rollout, run in parallel, and grade from its own database, so the verifier that filtered the SFT data returns a grounded, verifiable reward rather than one inferred from pixels. The same worlds that benchmark an agent train one.

Figure 12: Reinforcement learning on an Echoverse RLE. From the SFT policy we roll out a group of trajectories in one world; each is a sequence of act and execute steps that changes the database. A grader, the same grounded verifier that filtered the SFT data, sits outside the environment and scores each rollout’s final database state into a reward. The group of rewards updates the policy, and the loop repeats across every training world.

We take the SFT model as the starting policy and run RL against five worlds: EchoBank, EchoForge, EchoForum, EchoStay, and EchoTunes. Tasks come from the harder end of each world, where the SFT policy still leaves headroom, and each update draws on several graded rollouts. Each rollout earns two rewards: a trajectory reward from our database-grounded verifier (LLM judge GPT-4.1), and a dense per-step reward from a multimodal judge that grades each screenshot (GPT-4.1 vision). We train on roughly 100 tasks per world beyond the SFT data, for two epochs. On a held-out set of 25 tasks per world, the judged score rises from 58% to 69%. The teacher taught it what to do; the world taught it when to stop, when to recover, and when to give up.

Figure 13: Reinforcement learning on five worlds, over twoepochs. Left: the held-out judge score (25 tasks per world) climbs from 58% to 69%. Right: the critic’s mean score, the RL reward signal, trends up through training. The reward sums a trajectory reward from our database-grounded verifier (LLM judge GPT-4.1) and a dense per-step reward from a multimodal judge (GPT-4.1 vision). Where this leaves us

A world is no longer a fixed benchmark you score against; it is a training surface you keep improving, where the same graded run that measures the model also sharpens the world that judged it. Deep worlds transferred where shallow clones pulled capability down; one widget rebuilt in a hundred forms taught a skill that reached the live web; co-evolution moved both sides at once; and reinforcement against the same worlds pushed the agent past imitation, lifting held-out performance and trimming wasted steps.

The durable advantage is not the largest inventory of synthetic websites. It is a factory that diagnoses what an agent cannot yet do, builds or repairs the world that teaches it, protects the capability already earned, and runs the loop again. The next turns scale three fronts at once. First, more deep worlds for the closed domains public benchmarks cannot reach. Second, more capability worlds for the interactions models keep failing. And, above all, more reinforcement against those grounded worlds: longer runs, harder tasks, and wider reward exploration that push the agent’s behavior and its performance further than imitation ever could. The levers compound: deeper and broader worlds make stronger RL, stronger RL boosts the agent, and every round exposes the next capability to build.

We are releasing a piece of the factory: environment code and graded test tasks for four worlds, two deep domains (EchoStay and EchoForge) and two capability worlds (the datepicker and nested-filter, each with an in-distribution and a held-out split). Every task carries the database-grounded verifier that scores it, so the same worlds can benchmark an agent or train one. Code and tasks: https://aka.ms/echoverse

When worlds grow at the frontier of an agent’s competence, evaluation stops being a scoreboard and becomes the engine that decides what to build next: worlds that keep learning alongside the agents they train.

Acknowledgments

We thank Alexey Taymanov, Andrew Zhao, Aravind Rajeswaran, Corby Rosset, Hussein Mozannar, Luiz Do Valle, Sara Abdali, Spencer Whitehead, Vibhav Vineet, Zach Nussbaum, Yadong Lu, Pashmina Cameron, Rafah Hosn, and Chinmay Karkar for their valuable help, insightful discussions, and continued support throughout this work.

Opens in a new tab

The post Echoverse: Deep, evolving environments for computer-use agents appeared first on Microsoft Research.

Categories: Microsoft

EvoLib: Turning experience into evolving knowledge

Microsoft Research - Thu, 07/30/2026 - 18:00
At a glance
  • Self-supervised. EvoLib enables large language models to learn from their own experience during inference, without requiring ground-truth labels or external feedback.
  • From experience to knowledge. EvoLib transforms past attempts into reusable skills and reflective insights that can be applied to future tasks.
  • Knowledge that evolves. Useful skills and insights are continually refined, consolidated, and reweighted, turning instance-specific observations into increasingly general knowledge over time.
  • Learning that transfers across tasks. By turning experience into reusable knowledge, EvoLib helps AI models learn from past successes and failures and evolve the knowledge that has the highest potential on improving future performance.
  • Built for today’s AI models. As EvoLib does not require model updates, it can be applied to any black-box language models and AI systems deployed through APIs.

Memory has become an important AI agent capability: the ability to store and retrieve past experiences. But memory alone is not learning. A collection of past conversations, reasoning traces, or action histories can quickly grow into a vast archive of experiences, making it difficult to identify the most relevant knowledge for a new task—let alone refine and evolve this knowledge to improve performance over time.

Humans learn differently. We do not remember every detail of our past experiences. Instead, we remember what matters: strategies that work, mistakes to avoid, and skills that transfer across situations. Over time, these lessons are refined into increasingly general and reusable knowledge. This ability to transform experience into transferable, evolving knowledge is one of the foundations of human learning.

In our recent paper, Test-Time Learning with an Evolving Library, we explore how AI systems can learn from experience in a similar way. We introduce EvoLib, a framework that transforms raw experience into an evolving library of knowledge. Rather than treating memory as a growing archive of past experiences, EvoLib extracts reusable knowledge from those experiences and continually refines it as new experiences arrive. Through the evolution of library, skills become more general, insights become more accurate, and downstream performance gets improved consistently over time. In this way, AI agents can continually learn from accumulating experience without updating the underlying model.

How EvoLib Works

Unlike traditional AI memory systems that store raw experiences as static information, EvoLib is built around the idea of evolving knowledge. In EvoLib, a unit of knowledge can take the form of a reusable skill distilled from a successful solution or a reflective insight learned from mistakes. Rather than simply accumulating more memories over time, EvoLib continually refines, consolidates and reweights existing knowledge as new experiences arrive. Concretely, we design the following mechanisms around knowledge evolution:

  • Consolidation. As new knowledge is extracted from recent experience, EvoLib retrieves similar knowledge from the library and tries to consolidate it with the new knowledge into a more general and reusable one. This allows knowledge to move beyond individual experiences and become applicable across tasks.
  • Weighting mechanism. EvoLib continually updates the importance of each knowledge unit based not only on its immediate utility on the current task, but also on how much it contributes to generating useful knowledge on future tasks. Over time, knowledge with the greatest long-term impact naturally becomes more prominent in the library.
Figure 1. EvoLib transforms raw experiences into reusable skills and insights, then continually evolves them through consolidation and dynamic weighting. Key Results

To evaluate EvoLib, we tested it across a diverse set of challenging tasks with different types of experiences and demands for learning:

  • Solving mathematical reasoning problems
  • Writing code to perform the given tasks under efficiency constraints
  • Making decisions to explore and interact with an environment to perform long-horizon tasks

Across these tasks, EvoLib consistently outperforms the top retrieval-based memory approaches and other abstract memory mechanisms with more efficient token usage.

We also evaluated how effectively EvoLib converts test-time compute into performance gains through continually evolving knowledge. Figure 2 compares EvoLib against both compute scaling methods that perform each task in isolation and strong memory-based learning approaches. Each curve shows how performance improves as the amount of test-time compute increases.

Across all three benchmarks, EvoLib achieves higher performance throughout most of the compute range and improves performance more rapidly with increasing compute.

These results suggest that the key to better learning may not simply be storing more memories or spending more compute. Instead, the greatest gains come from transforming experience into reusable knowledge that can be continually refined and applied across tasks.

Figure 2. Across all tasks, EvoLib converts test-time compute into performance gains more effectively than existing methods. Robustness to random task order

A natural question is whether such learning depends heavily on the order in which tasks are encountered. In the real world, an AI system may face diverse types of tasks in arbitrary order, and a useful learning framework should be robust to the randomness in task order. To evaluate this, we measured the task performance on the same set of heterogeneous tasks but with different task orders. We found that EvoLib consistently improves over existing memory-based learning approaches and maintains stable performance across different orderings. This indicates that EvoLib can continually learn from diverse tasks even when they are interleaved, suggesting its practical advantage in real-world scenarios where an agent must handle and learn from a mixed stream of heterogeneous user requests without relying on a structured curriculum.

As AI systems take on longer-running and more complex tasks, learning from experience will become increasingly important. The future of AI may depend not only on larger models and more computation, but also on mechanisms that allow systems to continually accumulate, refine, and reuse knowledge.

EvoLib is one step toward that vision. By transforming experience into evolving knowledge, it enables AI systems to continually improve and adapt after deployment. Rather than repeatedly starting from scratch, future AI systems may be able to build upon an evolving library of reusable skills and insights, much like humans do.

Code and experiment results are available on GitHub (opens in new tab) to support future research on memory and knowledge evolution in AI systems.

Opens in a new tab

The post EvoLib: Turning experience into evolving knowledge appeared first on Microsoft Research.

Categories: Microsoft

Verifying Rust cryptography in SymCrypt, from standards to code

Microsoft Research - Mon, 07/13/2026 - 18:00
How Rust, Lean, Aeneas, and AI agents are helping scale formal verification for production cryptographic algorithms At a glance
  • SymCrypt develops new verified cryptography using Rust, Aeneas, and Lean to provide higher security assurance.
  • We prove that their code safely and correctly implements standard algorithms, notably for post-quantum cryptography.
  • We are releasing verified code, specs, properties, and proofs initially for SHA-3 and ML-KEM. 
  • Aeneas allows verifying a large subset of Rust code and provides efficient automation in Lean to support the proof effort.
  • Agents allow scaling automation by writing proofs that are independently-verifiable.
Introduction and motivation for formal verification

Cryptographic code sits at the foundation of modern computing. It protects operating systems, cloud services, firmware, messaging systems, and the protocols that connect them. Small mistakes can have outsized consequences: a single arithmetic slip, missing bounds check, or incorrect state transition can undermine the security of an otherwise sound design.

Testing and auditing remain essential, but they are not enough on their own. Cryptographic implementations are often optimized, constant-time, architecture-specific, and deliberately low level. The code that ships rarely looks like the clean algorithm in a standard: it contains reductions, bit manipulations, SIMD intrinsics, carefully shaped loops, and portability layers for many environments.

Formal verification addresses this gap by deploying machine-checked proofs instead of relying on testing alone. Rather than merely checking that the code usually behaves correctly, verification implements a precise mathematical specification for all inputs that satisfy the stated preconditions.

In June last year, Microsoft announced we would formally verify new algorithms written in Rust in SymCrypt, the cryptographic provider used across products and services including Windows and Azure. New cryptographic implementations are being written in safe Rust, then verified in the Lean (opens in new tab) formal proof framework using the Aeneas (opens in new tab) toolchain. This applies in particular to post-quantum cryptography, which require fast secure implementations of complex algorithms. This combination gives us two layers of assurance: Rust rules out broad classes of memory-safety bugs, while Lean proofs establish functional correctness against formal specifications derived from standards.

The result is a new verification methodology for production cryptography: verify code as developers write it, preserve performance-oriented implementation choices, and make the proof process scalable enough to keep up with an evolving codebase.

Figure 1. Agents (stochastic, in blue) and tools (algorithmic, in green) for software verification. Human effort focuses on reviewing formalization of standards and main properties. Agents write proofs and intermediate properties. Compilation, code extraction, and proof verification are deterministic, not agentic. Status of verification in SymCrypt

We have open sourced a SymCrypt branch (opens in new tab) that includes formal specifications and proofs. This public branch makes the proof artifacts available alongside the Rust algorithm implementations they validate, showing how the methodology applies to production cryptographic code. SymCrypt is not a standalone research prototype; it is Microsoft’s open-source cryptographic library used across products and services including Windows and Azure Linux.

This first release includes complete proofs for the Rust ML-KEM and SHA3 code that is being used in insiders builds of Windows today. SymCrypt is extending the same Rust, Lean, and Aeneas-based workflow to more Rust-native algorithms and integrating them into production versions for Windows and Linux, including for instance verified Rust code for, e.g., AES-GCM, FrodoKEM, and ML-DSA. The rest of this post uses this SymCrypt work as a concrete example, starting with how public standards become executable Lean specifications.

Turning standards into formal Lean specifications

The first step is to formalize what the algorithm is supposed to do. For cryptographic primitives, the source of truth is usually a public standard: a NIST specification, an IETF RFC, or another carefully reviewed algorithm description.

In our approach, the Lean specification is designed to stay close to the standard. When the standard describes a loop, an array update, or a mathematical operation, the Lean model follows the same structure wherever possible. This syntactic proximity matters: it makes the formal specification easier to audit because reviewers can compare the standard and the Lean side by side.

Lean also lets us write executable specifications. That means we can run the formal model against official test vectors to catch transcription errors, off-by-one mistakes, or misunderstandings of the standard. For algorithms such as ML-KEM, we can go further and prove high-level mathematical properties, such as showing that the formal model of the number-theoretic transform corresponds to the intended operation over the relevant polynomial ring.

A representative example is the number-theoretic transform (NTT) from ML-KEM. The standard describes the algorithm as an in-place transformation over 256 coefficients modulo q, with three nested loops that update pairs of coefficients using successive powers of the constant ζ (= 17).

Here is a direct translation of the NIST standard in Lean, trying to stick as close as possible to the original syntax:

The Lean version deliberately mirrors the structure of the standard: the same loop nest, the same zeta selection, and the same coefficient updates, allowing easy line-by-line human review. At the same time, it is executable and uses mathematical types, so it can be tested against known vectors and connected to higher-level theorems about the NTT’s algebraic meaning. In summary, the Lean specification is a concise, executable, mathematically meaningful model that tracks the standard closely enough to be reviewed by cryptographers and proof engineers alike.

Connecting the formal specification to the code

Once the specification is formalized, the next challenge is to connect it to the implementation. We do not ask developers to rewrite production cryptographic code in a verification-oriented language, nor do we generate code that product teams must then own. Instead, we verify the Rust code that engineers write, exactly as they write it.

Aeneas makes this possible by translating Rust’s mid-level representation into a pure Lean model. Rust’s ownership and borrowing discipline are crucial here. They let Aeneas safely eliminate much of the reasoning about pointer aliasing, liveness, and mutation that makes verification of C-style code so expensive.

For example, a Rust function that updates an array in place becomes, in Lean, a function that explicitly takes and returns a functional array. Mutable borrows are translated into value transformations. This preserves the behaviour that matters while presenting proof engineers with a functional model that is far easier to reason about.

Once in Lean, the function can be equipped with a theorem that states that it refines a formal specification. In other words, for every input satisfying the required bounds and well-formedness conditions, the implementation function returns the same mathematical result as the standard-derived Lean specification.

This style keeps responsibilities cleanly separated. Software engineers continue to write idiomatic, performant Rust. Verification engineers work against generated Lean models and prove theorems about them. The Rust code and the proofs live side by side, but the proof burden does not shape the code into something unnatural.

Going back to the NTT example, its Rust implementation is a function fn ntt(&mut [u16; 256]) that uses a mutable borrow to update an array in-place. The Lean translation purifies it into a function ntt : Array U16 256#usize → Result (Array U16 256#usize) that directly outputs the updated array, while wrapping it into a Result type to explicitly capture the fact that Rust functions may panic.

In this case, the theorem states that, if the array satisfies a well-formedness invariant (ensuring it represents a valid polynomial), then running the Rust model ntt returns the well-formed representation of the result of the mathematical specification Spec.ntt, modulo conversion from low-level arrays to high-level polynomials.

Scaling this to every function in real cryptographic code required substantial automation. Lean’s extensibility lets us build a gradient of automation with tactics for symbolic execution, arithmetic, arrays, and bit-vector reasoning. The experience becomes closer to debugging: automation handles the routine proof obligations, while engineers can inspect and refine the proof when a goal does not close automatically.

Supporting intrinsics and multiple architectures

Production cryptography cannot ignore hardware. SymCrypt must run across environments ranging from embedded and kernel contexts to cloud services. It also needs to take advantage of platform-specific instructions when they are available, including SIMD intrinsics and architecture-specific optimized paths.

A verification story that only works for a portable reference implementation is therefore incomplete. We need to verify the code that actually ships: dispatch logic, optimized routines, and target-specific variants included.

The code below is adapted from the ntt_layer  function that is internally used by the NTT. This function is compiled differently for x86-64 and aarch64, allowing dynamic dispatch to target-specific or portable implementations. On x86-64, it checks the availability of SSE2 instructions, while on aarch64 it checks for Neon.

.gist .gist-meta { display: none !important; }

As rustc’s output is inherently target specific, our toolchain compiles the code several times, one per compilation target for which verification is required, before merging the corresponding models. In effect, this merge operation turns the static dispatch permitted by the cfg attributes in the Rust code into a first layer of dynamic dispatch between x86-64 and aarch64 in the Lean model. Following what the Rust code does, these target specific models then themselves dynamically dispatch to the models of the XMM, Neon, and generic implementations.

Intrinsics require a slightly different treatment. Some low-level wrappers, especially those that manipulate raw pointers or expose platform instructions, are modelled by small, carefully reviewed Lean specifications. Others can be modelled using Rust code, which can be tested against hardware reference documentation, then translated and verified. The surrounding safe Rust code is then verified against those models. This keeps the trusted surface narrow while preserving the performance benefits of hardware acceleration.

The important point is that verification does not require giving up optimization. The methodology is designed to preserve the complexities of production code – including intrinsics, dispatch, and platform-specific implementations – while still proving a single, auditable correctness statement.

Reflecting formal guarantees to the code developer

Formal verification only scales in an engineering organization if developers can understand what has been proved. It is not enough for a proof to exist in a repository; the guarantee must be visible, reviewable, and synchronized to the code that engineers maintain.

To support this, we expose verification results through automatically generated dashboards. These dashboards summarize theorems in developer-facing terms: preconditions, postconditions, covered functions, trusted models, and remaining assumptions. Engineers do not need to open Lean to see what has been verified. For instance, below is the page displayed by the dashboard for our ntt function.

Figure 2. Dashboard page for the theorem that shows the Rust function mlkem.ntt correctly implements the NTT specified in the NIST standard.

The specification clearly presents the theorem statement included in the Lean formal development: it separates the function input and preconditions from the post-condition by putting them above a horizontal line, and use fully qualified names with links to navigate to Rust and Lean definitions.

This feedback loop is especially useful for reviewing assumptions around intrinsics, target-specific code, and boundary conditions. A cryptographic developer can for example check whether the theorem fully captures what they expect their code to guarantee, and notice a formal statement is too weak, or a precondition is wrong.

The dashboards also aligns verification with continuous development. As Rust code changes, Lean models and proofs can be regenerated and replayed. When a proof breaks, that failure becomes a signal: either the implementation changed in a way that needs a proof update, or the change has exposed a real discrepancy with the specification.

This turns formal verification from a one-time research artifact into part of the engineering workflow.

Agentic proofs

The final ingredient is automation beyond traditional tactics: AI agents. Lean is well suited to this because proofs are machine-checked by a small trusted kernel. An agent may propose a proof script, but Lean independently verifies whether the proof is valid.

We use agents in two places. First, they help translate standards into Lean specifications. Because the resulting specification is executable, aligned to the original standard, tested against official vectors, supported by mathematical theorems, and much simpler than an implementation, it can be thoroughly audited even when an agent helped draft it.

Second, agents help write and maintain proofs. With the right libraries, tactics, examples, and documentation, agents can handle large amounts of proof work: unfolding generated models, applying specifications for helper functions, discharging arithmetic obligations, and repairing proofs after refactors.

This is particularly powerful because the Rust code and Lean proofs are separated. Agents do not need to annotate or modify the production Rust implementation to make a proof go through. They operate on the proof side, and the result is accepted only if Lean validates it and the final theorem states the desired guarantee without introducing unreviewed assumptions.

In practice, this changes the economics of verification. Work that previously required months of specialist effort can be accelerated dramatically. The proof engineer’s role shifts from writing every proof by hand to designing specifications, curating automation, reviewing theorem statements, and steering agents to complete their proofs.

Conclusion

Verified cryptography has often faced a difficult trade-off: the strongest guarantees came from specialized toolchains, generated code, and workflows that were hard for product teams to adopt. Rust, Lean, Aeneas, and agentic proof automation let us revisit that tradeoff.

By verifying Rust as written, deriving auditable specifications from standards, supporting optimized multi-architecture implementations, and reflecting proof results back to developers, formal verification can become part of normal cryptographic engineering rather than an after-the-fact research exercise.

That is the long-term promise: cryptographic code that remains fast, portable, maintainable, and developer-owned, while carrying machine-checked evidence that it implements the standards it is meant to realize.

Opens in a new tab

The post Verifying Rust cryptography in SymCrypt, from standards to code appeared first on Microsoft Research.

Categories: Microsoft

Verifying Rust cryptography in SymCrypt, from standards to code

Microsoft Research - Mon, 07/13/2026 - 18:00
How Rust, Lean, Aeneas, and AI agents are helping scale formal verification for production cryptographic algorithms At a glance
  • SymCrypt develops new verified cryptography using Rust, Aeneas, and Lean to provide higher security assurance.
  • We prove that their code safely and correctly implements standard algorithms, notably for post-quantum cryptography.
  • We are releasing verified code, specs, properties, and proofs initially for SHA-3 and ML-KEM. 
  • Aeneas allows verifying a large subset of Rust code and provides efficient automation in Lean to support the proof effort.
  • Agents allow scaling automation by writing proofs that are independently-verifiable.
Introduction and motivation for formal verification

Cryptographic code sits at the foundation of modern computing. It protects operating systems, cloud services, firmware, messaging systems, and the protocols that connect them. Small mistakes can have outsized consequences: a single arithmetic slip, missing bounds check, or incorrect state transition can undermine the security of an otherwise sound design.

Testing and auditing remain essential, but they are not enough on their own. Cryptographic implementations are often optimized, constant-time, architecture-specific, and deliberately low level. The code that ships rarely looks like the clean algorithm in a standard: it contains reductions, bit manipulations, SIMD intrinsics, carefully shaped loops, and portability layers for many environments.

Formal verification addresses this gap by deploying machine-checked proofs instead of relying on testing alone. Rather than merely checking that the code usually behaves correctly, verification implements a precise mathematical specification for all inputs that satisfy the stated preconditions.

In June last year, Microsoft announced we would formally verify new algorithms written in Rust in SymCrypt, the cryptographic provider used across products and services including Windows and Azure. New cryptographic implementations are being written in safe Rust, then verified in the Lean (opens in new tab) formal proof framework using the Aeneas (opens in new tab) toolchain. This applies in particular to post-quantum cryptography, which require fast secure implementations of complex algorithms. This combination gives us two layers of assurance: Rust rules out broad classes of memory-safety bugs, while Lean proofs establish functional correctness against formal specifications derived from standards.

The result is a new verification methodology for production cryptography: verify code as developers write it, preserve performance-oriented implementation choices, and make the proof process scalable enough to keep up with an evolving codebase.

Figure 1. Agents (stochastic, in blue) and tools (algorithmic, in green) for software verification. Human effort focuses on reviewing formalization of standards and main properties. Agents write proofs and intermediate properties. Compilation, code extraction, and proof verification are deterministic, not agentic. Status of verification in SymCrypt

We have open sourced a SymCrypt branch (opens in new tab) that includes formal specifications and proofs. This public branch makes the proof artifacts available alongside the Rust algorithm implementations they validate, showing how the methodology applies to production cryptographic code. SymCrypt is not a standalone research prototype; it is Microsoft’s open-source cryptographic library used across products and services including Windows and Azure Linux.

This first release includes complete proofs for the Rust ML-KEM and SHA3 code that is being used in insiders builds of Windows today. SymCrypt is extending the same Rust, Lean, and Aeneas-based workflow to more Rust-native algorithms and integrating them into production versions for Windows and Linux, including for instance verified Rust code for, e.g., AES-GCM, FrodoKEM, and ML-DSA. The rest of this post uses this SymCrypt work as a concrete example, starting with how public standards become executable Lean specifications.

Turning standards into formal Lean specifications

The first step is to formalize what the algorithm is supposed to do. For cryptographic primitives, the source of truth is usually a public standard: a NIST specification, an IETF RFC, or another carefully reviewed algorithm description.

In our approach, the Lean specification is designed to stay close to the standard. When the standard describes a loop, an array update, or a mathematical operation, the Lean model follows the same structure wherever possible. This syntactic proximity matters: it makes the formal specification easier to audit because reviewers can compare the standard and the Lean side by side.

Lean also lets us write executable specifications. That means we can run the formal model against official test vectors to catch transcription errors, off-by-one mistakes, or misunderstandings of the standard. For algorithms such as ML-KEM, we can go further and prove high-level mathematical properties, such as showing that the formal model of the number-theoretic transform corresponds to the intended operation over the relevant polynomial ring.

A representative example is the number-theoretic transform (NTT) from ML-KEM. The standard describes the algorithm as an in-place transformation over 256 coefficients modulo q, with three nested loops that update pairs of coefficients using successive powers of the constant ζ (= 17).

Here is a direct translation of the NIST standard in Lean, trying to stick as close as possible to the original syntax:

The Lean version deliberately mirrors the structure of the standard: the same loop nest, the same zeta selection, and the same coefficient updates, allowing easy line-by-line human review. At the same time, it is executable and uses mathematical types, so it can be tested against known vectors and connected to higher-level theorems about the NTT’s algebraic meaning. In summary, the Lean specification is a concise, executable, mathematically meaningful model that tracks the standard closely enough to be reviewed by cryptographers and proof engineers alike.

Connecting the formal specification to the code

Once the specification is formalized, the next challenge is to connect it to the implementation. We do not ask developers to rewrite production cryptographic code in a verification-oriented language, nor do we generate code that product teams must then own. Instead, we verify the Rust code that engineers write, exactly as they write it.

Aeneas makes this possible by translating Rust’s mid-level representation into a pure Lean model. Rust’s ownership and borrowing discipline are crucial here. They let Aeneas safely eliminate much of the reasoning about pointer aliasing, liveness, and mutation that makes verification of C-style code so expensive.

For example, a Rust function that updates an array in place becomes, in Lean, a function that explicitly takes and returns a functional array. Mutable borrows are translated into value transformations. This preserves the behaviour that matters while presenting proof engineers with a functional model that is far easier to reason about.

Once in Lean, the function can be equipped with a theorem that states that it refines a formal specification. In other words, for every input satisfying the required bounds and well-formedness conditions, the implementation function returns the same mathematical result as the standard-derived Lean specification.

This style keeps responsibilities cleanly separated. Software engineers continue to write idiomatic, performant Rust. Verification engineers work against generated Lean models and prove theorems about them. The Rust code and the proofs live side by side, but the proof burden does not shape the code into something unnatural.

Going back to the NTT example, its Rust implementation is a function fn ntt(&mut [u16; 256]) that uses a mutable borrow to update an array in-place. The Lean translation purifies it into a function ntt : Array U16 256#usize → Result (Array U16 256#usize) that directly outputs the updated array, while wrapping it into a Result type to explicitly capture the fact that Rust functions may panic.

In this case, the theorem states that, if the array satisfies a well-formedness invariant (ensuring it represents a valid polynomial), then running the Rust model ntt returns the well-formed representation of the result of the mathematical specification Spec.ntt, modulo conversion from low-level arrays to high-level polynomials.

Scaling this to every function in real cryptographic code required substantial automation. Lean’s extensibility lets us build a gradient of automation with tactics for symbolic execution, arithmetic, arrays, and bit-vector reasoning. The experience becomes closer to debugging: automation handles the routine proof obligations, while engineers can inspect and refine the proof when a goal does not close automatically.

Supporting intrinsics and multiple architectures

Production cryptography cannot ignore hardware. SymCrypt must run across environments ranging from embedded and kernel contexts to cloud services. It also needs to take advantage of platform-specific instructions when they are available, including SIMD intrinsics and architecture-specific optimized paths.

A verification story that only works for a portable reference implementation is therefore incomplete. We need to verify the code that actually ships: dispatch logic, optimized routines, and target-specific variants included.

The code below is adapted from the ntt_layer  function that is internally used by the NTT. This function is compiled differently for x86-64 and aarch64, allowing dynamic dispatch to target-specific or portable implementations. On x86-64, it checks the availability of SSE2 instructions, while on aarch64 it checks for Neon.

.gist .gist-meta { display: none !important; }

As rustc’s output is inherently target specific, our toolchain compiles the code several times, one per compilation target for which verification is required, before merging the corresponding models. In effect, this merge operation turns the static dispatch permitted by the cfg attributes in the Rust code into a first layer of dynamic dispatch between x86-64 and aarch64 in the Lean model. Following what the Rust code does, these target specific models then themselves dynamically dispatch to the models of the XMM, Neon, and generic implementations.

Intrinsics require a slightly different treatment. Some low-level wrappers, especially those that manipulate raw pointers or expose platform instructions, are modelled by small, carefully reviewed Lean specifications. Others can be modelled using Rust code, which can be tested against hardware reference documentation, then translated and verified. The surrounding safe Rust code is then verified against those models. This keeps the trusted surface narrow while preserving the performance benefits of hardware acceleration.

The important point is that verification does not require giving up optimization. The methodology is designed to preserve the complexities of production code – including intrinsics, dispatch, and platform-specific implementations – while still proving a single, auditable correctness statement.

Reflecting formal guarantees to the code developer

Formal verification only scales in an engineering organization if developers can understand what has been proved. It is not enough for a proof to exist in a repository; the guarantee must be visible, reviewable, and synchronized to the code that engineers maintain.

To support this, we expose verification results through automatically generated dashboards. These dashboards summarize theorems in developer-facing terms: preconditions, postconditions, covered functions, trusted models, and remaining assumptions. Engineers do not need to open Lean to see what has been verified. For instance, below is the page displayed by the dashboard for our ntt function.

Figure 2. Dashboard page for the theorem that shows the Rust function mlkem.ntt correctly implements the NTT specified in the NIST standard.

The specification clearly presents the theorem statement included in the Lean formal development: it separates the function input and preconditions from the post-condition by putting them above a horizontal line, and use fully qualified names with links to navigate to Rust and Lean definitions.

This feedback loop is especially useful for reviewing assumptions around intrinsics, target-specific code, and boundary conditions. A cryptographic developer can for example check whether the theorem fully captures what they expect their code to guarantee, and notice a formal statement is too weak, or a precondition is wrong.

The dashboards also aligns verification with continuous development. As Rust code changes, Lean models and proofs can be regenerated and replayed. When a proof breaks, that failure becomes a signal: either the implementation changed in a way that needs a proof update, or the change has exposed a real discrepancy with the specification.

This turns formal verification from a one-time research artifact into part of the engineering workflow.

Agentic proofs

The final ingredient is automation beyond traditional tactics: AI agents. Lean is well suited to this because proofs are machine-checked by a small trusted kernel. An agent may propose a proof script, but Lean independently verifies whether the proof is valid.

We use agents in two places. First, they help translate standards into Lean specifications. Because the resulting specification is executable, aligned to the original standard, tested against official vectors, supported by mathematical theorems, and much simpler than an implementation, it can be thoroughly audited even when an agent helped draft it.

Second, agents help write and maintain proofs. With the right libraries, tactics, examples, and documentation, agents can handle large amounts of proof work: unfolding generated models, applying specifications for helper functions, discharging arithmetic obligations, and repairing proofs after refactors.

This is particularly powerful because the Rust code and Lean proofs are separated. Agents do not need to annotate or modify the production Rust implementation to make a proof go through. They operate on the proof side, and the result is accepted only if Lean validates it and the final theorem states the desired guarantee without introducing unreviewed assumptions.

In practice, this changes the economics of verification. Work that previously required months of specialist effort can be accelerated dramatically. The proof engineer’s role shifts from writing every proof by hand to designing specifications, curating automation, reviewing theorem statements, and steering agents to complete their proofs.

Conclusion

Verified cryptography has often faced a difficult trade-off: the strongest guarantees came from specialized toolchains, generated code, and workflows that were hard for product teams to adopt. Rust, Lean, Aeneas, and agentic proof automation let us revisit that tradeoff.

By verifying Rust as written, deriving auditable specifications from standards, supporting optimized multi-architecture implementations, and reflecting proof results back to developers, formal verification can become part of normal cryptographic engineering rather than an after-the-fact research exercise.

That is the long-term promise: cryptographic code that remains fast, portable, maintainable, and developer-owned, while carrying machine-checked evidence that it implements the standards it is meant to realize.

Opens in a new tab

The post Verifying Rust cryptography in SymCrypt, from standards to code appeared first on Microsoft Research.

Categories: Microsoft

Aurora 1.5: Extending open foundation models for weather and Earth-system applications

Microsoft Research - Thu, 07/09/2026 - 18:46
At a glance
  • Aurora 1.5 is a major extension of Microsoft’s Aurora Earth System foundation model that adds 22 more weather variables relevant to energy, agriculture, transport, and climate risk, along with hourly temporal resolution and probabilistic ensemble forecasting.
  • Released as open source on GitHub with model checkpoints on Hugging Face, Aurora 1.5 enables researchers and developers to use, evaluate, and build on the model.
  • Aurora 1.5 connects open research to Microsoft Weather services, linking the model with data, infrastructure, managed access, and operational use for weather and Earth-system applications.

Aurora 1.5 is a major update to the open Aurora Earth-system foundation model, adding 22 new weather variables for a broader view of atmospheric conditions, hourly forecasts, and probabilistic ensemble forecasting. Developed by Microsoft Weather as an extension of the original model from Microsoft Research AI for Science, Aurora 1.5 shows how frontier research can move into broader use: open for researchers and developers to evaluate and extend, and designed to support customers where additional data, infrastructure, and operational assurance is needed. As climate and weather-related risks continue to affect communities, infrastructure, and economies worldwide, advances in Earth-system forecasting can help improve preparedness and decision-making.

What is Aurora?

Aurora is a foundation model for the Earth system developed by Microsoft Research AI for Science, first introduced in 2024 and published in Nature (opens in new tab) in 2025. It showed that a single model could be adapted to medium-range weather, ocean waves, atmospheric chemistry, and emerging climate applications, including high-resolution weather forecasting through fine-tuning. Its growing use has reinforced the value of an open, collaborative model that is easier to adapt, evaluate, and put to use. 

This next phase of Aurora (opens in new tab) builds on that foundation by making the model openly available for the global community to adapt, extend, and build on. 

What is new in Aurora 1.5?

Aurora 1.5 advances the broader effort to make open weather foundation models practical and scalable for organizations that rely on atmospheric and Earth-system intelligence. Alongside new variables and higher temporal resolution, Aurora 1.5 adds one of the most requested capabilities from users: ensemble forecasting. Because forecasts are sensitive to initial conditions and model uncertainty, ensembles run multiple simulations to show the range and likelihood of possible outcomes. Aurora 1.5 builds on Microsoft Research’s scientific foundation with new product engineering, cloud infrastructure, managed access, and decision-support capabilities. Together, these advances make Aurora 1.5 a valuable enterprise-grade weather solution for organizations. 

Figure 1: Illustration of the capabilities of Aurora 1.5 ensemble for predicting new impactful parameters such as total cloud cover and solar radiation. Ensemble mean and standard deviation are shown. 

The breadth update adds 22 new variables to Aurora’s original 4, including representative surface, pressure-level, wind, temperature, humidity, precipitation, and radiation fields. That broader coverage makes the model more relevant for sectors that depend on integrated Earth-system signals, from energy and agriculture to transport and resilience planning. 

The update to hourly temporal resolution enables fine-grained detail for precision operational guidance, such as the onset of precipitation, trade decisions, or a landfalling tropical cyclone. 

“Aurora 1.5 is a meaningful step toward making weather foundation models more open, useful, and practical. By releasing the model openly, we give researchers, developers, and organizations a clearer path to evaluate it, adapt it, and understand where it can help. Microsoft Weather’s role is to connect that open research foundation with the data, infrastructure, and applied workflows required by enterprises to use weather intelligence responsibly and with confidence.”

Sridhar Iyer, Corporate Vice President, Microsoft AI

Spotlight: Microsoft research newsletter

Microsoft Research Newsletter

Stay connected to the research community at Microsoft.

Subscribe today Opens in a new tab Ensemble Forecasting in Aurora 1.5 Unlocks More Confident Decisions in the Face of Weather Uncertainty

The ensemble version of Aurora 1.5 introduces stochastic perturbations to represent model uncertainty, allowing the generation of multiple forecast members to estimate the spread of possible futures. For a multitude of applications including power systems, transport, agriculture, extreme-weather planning, and climate risk, the model distribution matters as much as the best estimate. 

This ensemble capability was developed through multi-stage fine-tuning on top of the original Aurora model. After expanding the variable set and adding hourly temporal resolution, the team introduced controlled perturbations into the model’s latent conditioning pathway and optimized the ensemble for probabilistic forecast quality. A final round of auto-regressive fine-tuning on ECMWF High Resolution (HRES) analysis data from 2018 to 2023 improved rollout behavior and stability.

Figure 2. Comparing Aurora 1.5’s probabilistic forecasts with the ECMWF ensemble forecast. The shading shows relative probabilistic forecast error, using ECMWF ENS as the baseline: blue areas indicate where Aurora 1.5 performs better, and red areas indicate where it performs worse. Across upper-air geopotential, temperature, and humidity, together with five surface variables, Aurora 1.5 outperforms ECMWF ENS on 88.9% of the evaluated variable-and-lead-time targets. 

Aurora’s ensemble approach summarizes uncertainty across multiple model runs. Its probabilistic forecasts outperform those of the state-of-the-art ECWMF dynamical ensemble on 88.9% of evaluated targets (Figure 1). In evaluations on all 2024–2025 tropical cyclones, Aurora 1.5 substantially reduced track errors, including roughly one-third lower track error when comparing the ensemble median to the original Aurora. An example for the devastating Hurricane Helene shows how Aurora 1.5’s skill translates to high-impact weather applications. 

Figure 3. Hurricane Helene ensemble forecast from Aurora 1.5, showing multiple plausible storm tracks starting at 0 UTC on September 24, 2024. The probabilistic ensemble forecast envelops the verified track, effectively capturing uncertainty in the storm’s progression. Figure 4. Aurora 1.5 reduces track error relative to the original model across lead times. Ensemble mean and median tracks are used for diagnostics, with the median showing the strongest gains, reaching roughly one-third lower error by day 5. Results reflect track position only.  Beyond weather: Aurora as an Earth-system foundation

Beyond medium-range weather applications, Terradot – part of the Microsoft Climate Innovation Fund portfolio—is working with the AI for Good Lab (opens in new tab) and the Microsoft Research Accelerator on TerraNova, using Aurora-derived weather representations (opens in new tab) to estimate and optimize carbon dioxide removal from enhanced rock weathering under real field conditions. Sasankh Munukutla, Co-Founder of Terradot, highlights, “By building on Aurora, we’re significantly advancing our R&D timelines and accelerating our path towards gigaton-scale carbon removal.” This work shows how Earth-system foundation models can support climate mitigation and public-interest science beyond forecasting, including settings where rigorous evaluation and responsible deployment matter.

Aurora is also being explored with partners such as the UK Met Office, exploring how foundation models can work alongside established physics-based systems to tackle problems from weather to climate time scales. The aim is faster, more flexible forecasts that support decision-making without replacing the science behind trusted prediction. 

“Microsoft’s Aurora model is an exciting and promising tool, enabling Met Office scientists to bring their data and expertise to help solve climate problems and provide new kinds of climate information. Met Office and Microsoft scientists and engineers are working together every day to translate lessons from AI weather prediction into the climate information space, sharing expertise in data science and climate science. Aurora is a great platform for learning how to translate these tools for use in climate projection to make the AI climate models of the future.”

— Doug McNeall, Science lead for Data-Driven Climate Modelling, Met Office Hadley Centre  Connecting open models to operational use

Microsoft connects open research, product engineering, responsible deployment, and partner ecosystems so that models can move from scientific advance to evaluated operational use. As an example, Aurora began in Microsoft Research AI for Science and is now being built on for operational use by Microsoft Weather, with AI for Good helping to evaluate public-interest applications. The platform path brings Aurora into Microsoft Foundry and Planetary Computer Pro, alongside Agent skills and Azure services that connect models with geospatial data, scalable infrastructure, and applied workflows. BKW provides an early proof point: the company is using Aurora 1.5 alongside existing operational Microsoft Weather models to support energy operations where weather-dependent generation, infrastructure planning, and environmental data need to come together. 

“This collaboration demonstrates how advanced AI capabilities and robust cloud infrastructure can be applied to one of the most strategic domains — energy, where weather plays a fundamental role. In a time of accelerated transformation, it supports our ambition to operate increasingly renewable-based systems, where generation is inherently weather-dependent, and to better anticipate and manage this variability with greater confidence and precision.” 

Farhat Quiñones Yamshid, Lead, AI and Technology, BKW  From open research to broader impact

Aurora’s open-source availability is intended to help researchers, agencies, companies, and civil society evaluate, apply, and extend the model. Microsoft Weather is building on that open foundation to deliver easier access to Aurora forecasts through managed services, integrations, and responsible deployment paths for organizations that depend on weather and Earth-system intelligence.

Foundation models should complement—not replace—physics-based models and domain expertise. The opportunity is to use them responsibly, with careful evaluation and transparency, and to invite researchers, agencies, companies, and public-interest partners to test where Aurora and related Microsoft Weather capabilities can improve forecasting, planning, and climate resilience in their own settings.

About Microsoft Weather 

Microsoft Weather is the AI-based forecasting team behind weather experiences across Windows, Bing, Copilot, Edge, and MSN, reaching more than a billion devices across 180 countries. The team has been applying AI to operational weather forecasting for more than seven years and has built a proven track record of delivering high-quality forecasts at global scale. Microsoft Weather has won multiple forecasting competitions and was ranked the world’s most accurate global forecast provider by an independent third party for three consecutive years from 2022 to 2024. Building on today’s Aurora 1.5 announcement, the team plans to extend this work in the coming months with additional fit-for-purpose AI weather models designed for enterprise scenarios where forecast quality, speed, uncertainty, and operational decision support matter most.

If you are interested in exploring Aurora and Microsoft Weather solutions for commercial or organizational applications, please contact us at AIWeatherClimate@microsoft.com 

Aurora 1.5 on Microsoft Foundry Aurora 1.5 on GitHub Aurora 1.5 paper Agent Skills for adapting Aurora to new applications

Opens in a new tab

The post Aurora 1.5: Extending open foundation models for weather and Earth-system applications appeared first on Microsoft Research.

Categories: Microsoft

Flint: A visualization language for the AI era

Microsoft Research - Wed, 07/08/2026 - 18:00
At a glance
  • Polished charts from simple specs. Flint allows AI agents to reliably generate expressive, visually polished charts from simple, human-editable specifications.
  • Semantic types guide design. Flint leverages semantic data types to express meanings of data. They help the compiler choose appropriate scales, baselines, formatting, and color schemes.
  • Layouts adapt to the data. Flint automatically manages sizing, spacing, labels, and layout so charts remain readable as cardinality and density change, without explicit user configurations.
  • One spec can target multiple backends. A single Flint specification can compile to Vega-Lite, Apache ECharts, or Chart.js without rewriting the chart from scratch.
  • Built for agent workflows. The open-source project includes the flint-chart library and the flint-chart-mcp server, so agents can create, validate, and render charts directly in chat or coding environments.
Figure 1. Flint supports a diverse collection of visualizations with its simple spec, which can be rendered with visualization libraries like Vega-Lite, Echarts, and Chart.js.

Creating a good chart requires many design decisions: how dates should be parsed, whether a scale should start at zero, how values should be formatted, how much room labels need, and which colors make the data easier to read. Modern visualization libraries such as Vega-Lite, Apache ECharts, and Chart.js expose these controls, but there is a trade-off: Short specifications that rely on system defaults often produce uninspiring charts, while polished visualizations require detailed specifications with purposely chosen parameters that are often verbose, fragile, and error-prone.

This trade-off becomes sharper as large language models (LLMs) and AI agents take on more visualization work. Agents are especially prone to errors when they must manage complex, low-level specification details, and the resulting fragile code can be difficult for people to inspect, repair, or reuse. Ideally, we need something in between: a compact specification that agents can produce reliably, people can edit directly, and a system can compile into a well-designed chart.

To address this challenge, we introduce Flint (opens in new tab), a visualization intermediate language for AI-driven chart creation. Flint helps AI agents create expressive, attractive charts from simple, human-editable chart specs. Instead of requiring verbose low-level parameters for scales, axes, spacing, and layout, the Flint compiler derives optimized chart settings from the data, semantic types, chart type, and encodings. The same Flint spec can render through multiple backends, including Vega-Lite, Apache ECharts, and Chart.js.

Figure 2. Flint compiles a compact, human-editable chart specification into a complete backend-native specification and rendered visualization. In this heatmap example, the Flint spec names semantic types (period as YearMonth, newUsers as Profit) and maps fields to visual channels. The compiler derives the Vega-Lite details, including temporal parsing, axis formatting, color scale, cell sizing, legend configuration, and layout. How Flint works

Figure 2 illustrates the how the Flint compiler turns a compact chart specification into a refined heatmap.

To produce a high-quality heatmap, traditionally, we need to explicitly tell the system with low-level chart properties about how to process the period field, how to properly label MonthYear values, size individual heatmap cells, and choose a color scale that appropriately represents positive and negative newUsers values. Without these configurations, visualization libraries must guess from field names and raw values, which can lead to charts that are technically valid but potentially misleading. While they are important, hard-coding these details can be difficult and error-prone, and they make specification fragile and hard for users to understand or adapt.

In Flint, these low-level details are systematically managed, where the compiler infers them from high-level data and chart specifications. Here, the data specification captures semantic types and optional metadata, and the chart specification defines the chart type and maps fields to visual channels such as x, y, color, size, or facet. From this information, the compiler derives the parsing rules, scales, axes, aggregations, formatting, color schemes, layout, and generates the backend-native specification, which is used to render the final polished visualization. This frees users from explicitly setting fragile and error-prone low-level details.

Furthermore, because the intermediate representation is separate from any single rendering library, Flint can target backends with very different APIs and programming models. Users can keep the same compact chart intent while compiling to Vega-Lite, ECharts, or Chart.js, and choose the backend whose capabilities best fit the visualization.

Spotlight: AI-POWERED EXPERIENCE

Microsoft research copilot experience

Discover more about research at Microsoft through our AI-powered experience

Start now Opens in a new tab Flint for AI-assisted visualization

Flint is well suited to LLM-based chart generation because semantic types are often easier for models to infer than the full set of low-level visualization parameters. Field names, value patterns, and common data knowledge can help an agent recognize whether a column represents a date, price, percentage, country, ranking, or correlation. Once those meanings are explicit, the compiler can handle many design decisions that would otherwise appear as brittle, library-specific code.

In our research study, we compared Flint with DirectVL, a baseline that asks the model to directly generate full (more complex) Vega-Lite specifications in a LLM self-evaluation pipeline. Across three tested models based on testing data from Tidy Tuesdays, Flint received higher overall LLM-judge scores: 16.27 vs. 15.91 with GPT-5.1, 16.16 vs. 15.60 with GPT-5-mini, and 15.91 vs. 15.34 with GPT-4.1. In fact, Flint has been so powerful and reliable that it is now used to power Data Formulator (opens in new tab), a Microsoft Research project for AI-assisted data analysis and visualization.

To make Flint easy for your agents to access, we also release flint-chart-mcp, a Model Context Protocol (MCP) server that allows agents to create, validate, and render charts inside a chat or coding environment. MCP calls can embed data inline or read configured local files, and the server can open an interactive chart view so users can inspect and refine the results.

Figure 3. Once you set up the flint-chart-mcp with your favorite AI client, the agent can generate interactive visualizations powered by Flint to answer your data exploration questions. Try Flint

Flint is open source and ready to use:

Flint points toward a shared semantic layer for visualization, where people and AI agents can work with compact chart intent while a compiler handles the careful low-level details. We invite the community to explore the project and build on it.

Opens in a new tab

The post Flint: A visualization language for the AI era appeared first on Microsoft Research.

Categories: Microsoft

SkillOpt: Agent skills as trainable parameters

Microsoft Research - Tue, 06/30/2026 - 18:50
At a glance
  • AI agents often fail because their instructions, or skills, are manually modified with no guarantee of improvement. SkillOpt turns skill editing into a training process, making agent behavior more reliable without changing model weights.
  • SkillOpt treats an agent skill file as a trainable parameter outside a frozen target model, turning skill writing from one-shot prompting into a controlled optimization process.
  • Across six benchmarks, seven target models, and three execution modes, SkillOpt is the best or tied-best method in all 52 evaluation cells, improving performance without updating model weights.
  • SkillOpt keeps skills compact and auditable through bounded text edits, validation gating, rejected-edit feedback, and slow/meta updates, avoiding uncontrolled prompt drift.
  • The optimized skills transfer across model scales, agent harnesses, and related tasks, suggesting that they capture reusable workflow knowledge rather than benchmark-specific instructions.

Large language models (LLMs) are increasingly deployed as agents that gather evidence, call tools, and execute multi-step tasks. For these agents, the hard problem is no longer whether they can call a tool, but whether they can complete tasks reliably and consistently. Today, agent skills typically come from three sources: experts write them by hand, a frontier model generates them one-shot, or the agent loosely revises them after execution. None of these approaches behaves like a deep-learning optimizer. They lack step-size control, held-out validation, and any memory of revisions that failed. As a result, skills tend to grow longer and drift with each rewrite, and a revision that seems perfectly reasonable can quietly degrade real task performance. This uncontrolled skill evolution has become a major obstacle on the path from agent prototype to dependable, production-grade deployment.

In our recent paper, SkillOpt: Executive Strategy for Self-Evolving Agent Skills, we reframe the question from “how do we write a better prompt?” to “how do we train the skill?” SkillOpt treats the skill file as a trainable parameter living outside a frozen target model, bringing a training-style optimization loop, consistent gains across 52 evaluation cells, and a compact skill file that stays readable, auditable, and transferable.

Figure 1. A frozen target model executes tasks while a separate optimizer model trains the skill layer from trajectory feedback, exporting the reusable skill file best_ skill.md through validation gating. How SkillOpt works Video 1. SkillOpt’s optimization loop, from trajectory collection to the exported skill file.

SkillOpt organizes skill editing as a forward–backward–update cycle in text space. In the forward pass, the frozen target model executes a batch of training tasks with the current skill; the rollout batch size controls how much evidence each update receives. In the backward pass, a separate optimizer model reads the resulting trajectories in reflection minibatches, distilling patterns to preserve from successful trajectories and patterns to correct from failures.

In the update step, the optimizer proposes small add, delete, and replace edits; candidate edits are merged, deduplicated, ranked, and clipped by a textual learning rate—a per-step edit budget. Every candidate skill must then pass a strict validation gate: it is adopted only if it scores strictly higher than the current skill on the held-out validation split. Rejected edits are not discarded; they enter a rejected-edit buffer that serves as negative feedback for later optimizer calls in the same epoch. On a slower cadence, an epoch-wise slow/meta update consolidates longer-horizon lessons that single batches cannot reveal (Figure 2). Together, bounded edits, validation gating, and best-version selection keep skill optimization controllable and auditable, so the skill converges instead of drifting.

Figure 2. The SkillOpt pipeline: trajectory collection, minibatch reflection, bounded text updates, validation gating, and epoch-wise slow/meta updates jointly constrain skill training. Consistent gains across benchmarks, models, and execution modes

We evaluated SkillOpt across six benchmarks (SearchQA, SpreadsheetBench, OfficeQA, DocVQA, LiveMathematicianBench, and ALFWorld), seven target models from frontier-scale GPT-5.5 to the small open-weight Qwen3.5-4B, and three execution modes (direct chat, Codex, and Claude Code). Counting each combination as one evaluation cell, When measured against human-written skills, one-shot LLM skills, Trace2Skill, TextGrad, GEPA, and EvoSkill, SkillOpt delivered the best or tied for -best results on all 52 cells. These performance improvements are unusually large for a method that updates no model weights. With GPT-5.5 in direct chat, SkillOpt raises the six-benchmark average from 58.8 to 82.3, a +23.5-point absolute improvement—and +5.4 points above an oracle that picks the single best competing method per cell. The largest gains appear on procedural benchmarks: SpreadsheetBench rises from 41.8 to 80.7, OfficeQA from 33.1 to 72.1, and LiveMathematicianBench from 37.6 to 66.9. The same interface carries over to agentic loops, lifting GPT-5.5 by +24.8 points inside Codex and +19.1 inside Claude Code over no skill.

PODCAST SERIES

AI Testing and Evaluation: Learnings from Science and Industry

Discover how Microsoft is learning from other domains to advance evaluation and testing as a pillar of AI governance.

Listen now Opens in a new tab A small model plus a skill file

Approaching the next model tier SkillOpt also narrows the gap between small or open-weight models and frontier models—without changing any weights or adding any extra model calls at inference. After optimization, GPT-5.4-mini’s six-benchmark average (64.3) exceeds the no-skill baseline of the larger GPT-5.4 (59.7), and GPT-5.4-nano (57.4) exceeds the no-skill baseline of GPT-5.2 (51.3). Qwen3.5-4B, a 4-billion-parameter open-weight model, surpasses GPT-5.2’s no-skill baseline as well. Gains that once required a larger model can now be approximated by one optimized skill file.

Skills that transfer: train once, reuse everywhere

The optimized skill file captures reusable task-solving procedures rather than instructions overfit to a single model, benchmark, or execution environment. This is why the same skill can still improve performance when transferred across model scales, agent harnesses, and related tasks. In our transfer experiments, skills continued to deliver gains when moved across model scales, across execution harnesses, and to a nearby math benchmark. The clearest example is cross-harness transfer: a spreadsheet skill trained inside Codex, dropped into Claude Code with no further optimization, lifts the no-skill baseline from 22.1 to 81.8 (+59.7)—slightly above the 80.4 achieved by training directly inside Claude Code. Because the two harnesses expose different tool surfaces, this suggests SkillOpt learns general workflow logic, not just harness-specific recipes.

Compact, readable, and built from very few accepted edits

The deployed artifact, best_ skill.md , is neither an opaque parameter blob nor an ever-growing log. Across six case studies, the median final skill length is roughly 920 tokens, and because the validation gate rejects most proposals, only one to four edits are accepted into the final file. OfficeQA’s +39.0-point gain comes from a single accepted edit. The learned rules read like a seasoned practitioner’s advice. Component ablations confirm that the controls do the work: removing the rejected-edit buffer lowers scores on all three ablation benchmarks, and removing both the meta skill and the slow update drops SpreadsheetBench from 77.5 to 55.0. A new adaptation layer for the agent era SkillOpt points to a lighter-weight path for domain-adapting agents: instead of fine-tuning weights, hard-coding task logic, or hand-tuning prompts, teams can train a small, versionable, auditable natural-language skill layer—wherever automatic evaluation or a reliable verifier exists.

By bringing learning rates, schedules, validation splits, rejected samples, and slow updates to agent skills, SkillOpt suggests that training need not be limited to model weights. Procedural knowledge outside the model can also be optimized.

When that process is controlled, validated, and recorded, a natural-language skill becomes a stable, transferable, and reversible adapter between frontier-model capability and real-world workloads. Read the full paper, visit the project page at aka.ms/skillopt (opens in new tab), or explore the SkillOpt GitHub repository at github.com/microsoft/SkillOpt (opens in new tab). Teams building agentic workflows can use SkillOpt as a foundation for training reusable skills against their own tasks and verifiers. See also our companion project, SkillLens.

Paper GitHub SkillLens Project Page Opens in a new tab

The post SkillOpt: Agent skills as trainable parameters appeared first on Microsoft Research.

Categories: Microsoft
Syndicate content

eXTReMe Tracker