Blogroll

I turned this tiny $16 USB drive into the ultimate PC rescue tool

How-To Geek - 1 hour 1 min ago

If you're the person everybody calls when their PC breaks down, I'm sure you can relate to all those times when somebody randomly asks you to fix their PC on the spot.

Categories: IT General, Technology

Google TV Freeplay now has on-demand movies and shows, but it still trails rivals in one key category

How-To Geek - 1 hour 10 min ago

Google's answer to free streaming services like The Roku Channel and Tubi is now much more appealing if you like control over what you watch. Google TV Freeplay now has on-demand movies and TV shows as well as an expanded live channel lineup.

Categories: IT General, Technology

This American luxury SUV just became a much smarter buy

How-To Geek - 1 hour 16 min ago

For years, the Jeep Grand Wagoneer had a pretty obvious problem: it looked and felt like a luxury SUV, but it asked luxury-SUV money for a Jeep badge. That made it tough to justify next to established names like Cadillac and Lincoln.

Categories: IT General, Technology

Adding a teen to your car insurance may cost $300 extra a month—here's how to slash that

How-To Geek - 2 hours 16 min ago

Back-to-school season brings a list of new costs for families with a teen driver in the house, from a parking permit to gas money to maybe a vehicle of their own. Car insurance is yet another cost that joins that seemingly endless list.

Categories: IT General, Technology

6 BSDs worth trying instead of Linux

How-To Geek - 2 hours 21 min ago

When you think of open-source OSes, you might think of Linux distros, but BSD-based systems have been around for a long time. Here are some of the best BSD systems that you can try that can give Linux distros a run for their money.

Categories: IT General, Technology

The Anker Solix S2000 is on sale for under $650 — keep the refrigerator cooling for up to 35 hours

Mashable - 2 hours 45 min ago

SAVE $549.01: The Anker Solix S2000 portable power station is on sale at Amazon for $649.99, down from the list price of $1,199. That's a 46% discount.

Opens in a new window Credit: Anker Solix Anker Solix S2000 $649.99 at Amazon
$1,199 Save $549.01   Get Deal

When storms show up in the forecast, it's always smart to get prepared. That usually means making sure phones are charged, flashlights are available, and you have a few food options on hand. Speaking of food, if you're looking for a power source that'll take over keeping the refrigerator cool, check out this deal.

As of Aug. 11, the Anker Solix S2000 portable power station is on sale at Amazon for $649.99, down from the list price of $1,199. That's a 46% discount.

A portable power station can be useful in plenty of situations from camping to a power outage. Some models focus on portability while others are durable enough to leave out in the rain. If your main concern during a power outage is keeping the refrigerator on, the Anker Solix S2000 might be the perfect model.

Anker Solix designed the S2000 specifically for home use during a power outage. The design with AC ports on both the front and back means its great for setting on the counter and plugging the fridge into the back while keeping the front open for a lamp, the coffee maker, and recharging small gadgets.

SEE ALSO: I tried the Jackery FridgeGuard portable power station: It's both functional and pretty

With 2,010Wh of battery capacity, Anker says it can keep a fridge running for up to 35 hours. In part, that's thanks to Anker's low 6W power draw. With an under 10 millisecond uninterrupted power supply, you can keep the refrigerator plugged into the Solix and the Solix plugged into the wall on grid power. If the power cuts out, the Solix will take over in under 10 milliseconds.

If you're not looking forward to the upcoming storm season since that means worrying about food spoiling in the fridge during an outage, upgrade to the Solix S2000. It can keep foods cooling for over 24 hours and it's on sale for under $650.

Categories: IT General, Technology

Gemini pushed me to try Samsung Bixby, and it wins at one key thing

How-To Geek - 2 hours 46 min ago

Most people who switched away from Bixby didn’t leave because it was bad at what it was built to do. They left because something shinier showed up, and going back never crossed their mind. I was the same way. I moved from Bixby to Google Assistant to Gemini without really thinking about it, and I assumed Bixby had quietly died somewhere along the way because it was somehow bad. But it’s actually giving Gemini a run for its money.

Categories: IT General, Technology

Ford’s new AI assistant may make your owner’s manual obsolete

How-To Geek - 3 hours 16 sec ago

The days of digging through your car's owner's manual might soon be over. Ford is now broadly rolling out an AI assistant that can answer questions about your car, including some that would normally involve thumbing through a book or talking to a service technician.

Categories: IT General, Technology

Snag the beginner-friendly DJI Mini 3 drone for $50 off

Mashable - 3 hours 7 min ago

SAVE $50: As of August 11, get the DJI Mini 3 for $499 at Amazon, down from its usual price of $549. That's a discount of 9%.

Opens in a new window Credit: Amazon DJI Mini 3 $499 at Amazon
$549 Save $50   Get Deal

If you're interested in flying a drone, creating drone-centric content, or just taking some cool pictures, DJI is a brand that has plenty of options for you to choose from. While it's been difficult to obtain a DJI drone with the FCC ban, there are still occasionally models that go back in stock at Amazon. There's one on sale right now you can get for a great price, and if you're interested in grabbing it, you'll want to do so before it once again becomes unavailable.

As of August 11, get the DJI Mini 3 for $499 at Amazon, down from its usual price of $549. That's $50 off and a discount of 9%.

SEE ALSO: The ultimate travel vlogging setup: Save over $100 on the DJI Osmo Pocket 3 Vlog Bundle

This lightweight drone is just half a pound, which makes it an impressively small drone that's perfect for newbies. It can even fold up to become smaller, so if you want something compact and simple to use, this is the model to go after. But just because it's small, that doesn't mean it's bereft of features. It has 4K HDR video capabilities as well a gimbal design that allows it to shoot vertically with dynamic angles.

It offers about 38 minutes of flight time per battery charge, which you can extend up to 51 minutes if you pick up the Intelligent Flight Battery Plus as a separate purpose. Other than that, it also has built-in intelligent features of its own, like auto takeoff, precise hovering, and return to home capabilities so you don't have to do it all yourself.

If you're ready to try out drone flying, this is a great time to do so and a deal that's well worth starting with.

Categories: IT General, Technology

This family SUV feels like a huge upgrade without the huge price

How-To Geek - 3 hours 16 min ago

There are plenty of three-row SUVs that can haul a family, swallow a pile of luggage, and get everyone where they need to go. What’s harder to find is one that makes you feel like you spent considerably more than you actually did.

Categories: IT General, Technology

Take 200 bucks off the Ecovacs Deebot X11s Pro Omni and streamline your floor cleaning setup

Mashable - 3 hours 22 min ago

SAVE $200.99: As of Aug. 11, the Ecovacs Deebot X11S Pro Omni robot vacuum and mop is on sale at Amazon for $799 instead of its usual $999.99. That's a savings of 20%.

Opens in a new window Credit: Ecovacs Ecovacs Deebot X11S Pro Omni robot vacuum and mop $799 at Amazon
$999.99 Save $200.99   Get Deal

There's no denying that the robot vacuum market is severely oversaturated. With tons of brands, models, and features to navigate, we understand that choosing the right one isn't easy. This $200 discount on the Ecovacs Deebot X11s Pro Omni, however, might be the factor that finally helps you decide.

As of Aug. 11, the Ecovacs Deebot X11s Pro Omni robot vacuum mop combo is on sale at Amazon for only $799. It usually clocks in at $999.99, which means this deal saves you about 20% or $200.99. According to our favorite price-tracking tool, that's also its best price on record.

We have a few dealbreakers when it comes to buying a robot vacuum mop combo. The first is the lack of self-washing and drying the mop pads. Fortunately, the X11s Pro Omni tackles everything for you with its all-in-one Omni station. It handles all the maintenance, including heated high-pressure washing, auto-drying, auto dustbin emptying, water box auto-cleaning, and more.

Other dealbreakers include robot vacuums with less than 10,000 Pa suction power and the lack of obstacle avoidance and smart mapping on board. The Deebot X11s Pro Omni checks all the right boxes with up to 30,000 Pa suction power, advanced AI-powered navigation with smart obstacle avoidance, and smart control with YIKO voice assistance.

If you want a robot vacuum that will make your life a little easier and take both vacuuming and mopping off your hands, the Ecovacs Deebot X11s Pro Omni is a solid pick. And at 20% off, it's not too harsh on your wallet either.

Categories: IT General, Technology

Asus Gamer Day deals take hundreds off ROG laptops, desktops, and a mini PC

Mashable - 3 hours 46 min ago
Asus Gamer Days deals at a glance Best gaming laptop deal Asus ROG Strix G16 Gaming Laptop (Intel Core i9, 32GB RAM, 1TB) $1,799.99 (save $200) Get Deal Best gaming desktop deal Asus ROG G700 (Intel Core 7 265F, 32GB RAM, 1TB) $1,649.99 (save $150) Get Deal Best gaming mini PC deal Asus ROG NUC 16 (Intel® Core Ultra 9, 64GB, 1TB) $2,999 (save $200) Get Deal

Tech prices continue to rise, especially as the memory shortage gets more intense. That's not great news for anyone who uses tech, but it can be espeically painful for gamers. Running games tends to require a machine with higher-end specs and more memory. But if you've been looking at an Asus gaming set-up, there's good news in store.

The Intel Asus Gaming Days event is underway through Sept. 13. With it comes deals on gaming laptops, desktops, and mini PC options. The best deals take $200 off and come with the benefit of getting two free games. Eligible purchases will email a unique master key to claim your free games by Oct. 31.

Best gaming laptop deal Opens in a new window Credit: Asus Asus ROG Strix G16 Gaming Laptop (Intel Core i9, 32GB RAM, 1TB) $1,799.99 at Asus
$1,999.99 Save $200   Get Deal Why we like it

Built for speed, excellent visuals, and responsiveness, the Asus ROG Strix G16 is on sale for $1,799.99 for a $200 discount. This gaming laptop comes with the Intel Core i9 processor, 32GB memory, and 1TB SSD storage. In terms of display, you're in line for 165Hz FHD+ which is ideal for both gaming and streaming action content.

Asus also paid special attention to cooling with the ROG Strix G16, equipping it with an advanced cooling system that says quiet. Ports on the Strix G16 include an HDMI 2.1, Thunderbolt 4, 2.5G ethernet port, an audio jack, plus a few more.

Best gaming desktop deal Opens in a new window Credit: Asus Asus ROG G700 (Intel Core 7 265F, 32GB RAM, 1TB) $1,649.99 at Asus
$1,799.99 Save $150   Get Deal Why we like it

The 2025 Asus ROG G700 comes with the Intel Core Ultra 7 processor and NVIDIA GeForce RTX 5060Ti graphics card. The 32GB RAM and 1TB SSD make for quick multitasking and loading games. There's also an included RGB keyboard and mouse to complete your new gaming setup.

In terms of cooling, the Asus has three fans for intake and a rear exhaust. Plus, there's an included dust filter. Pair it with an excellent gaming monitor and you're in for the best winter gaming season yet.

Best gaming mini PC deal Opens in a new window Credit: Asus Asus ROG NUC 16 (Intel Core Ultra 9, 64GB, 1TB) $2,999 at Asus
$3,199 Save $200   Get Deal Why we like it

If you're in the market for a mini PC, the Asus ROG NUC 16 is included in Asus Gaming Days deals, taking $200 off the base price. Asus says the ROG NUC 16 is meant for hardcore gamers looking for advanced features, combining a compact size with powerful performance. Stand it up or set it down horizontally, and you're in line for great performance with an Intel Core Ultra 9 processor. For even better graphics, Asus used the NVIDIA DLSS 4.5 and the NVIDIA Reflex for low latency.

Thumb screws give you easy access for adding memory or storage without the need to track down tools. There's also a dedicated heat sink that Asus says reduces operating temperature to a cool 59 degrees.

Categories: IT General, Technology

I turned my old Samsung phone into the ultimate Fire TV replacement

How-To Geek - 3 hours 46 min ago

Many of the streaming media players on the market are built on Android—this includes Fire TV and, of course, Google TV. So, why can’t an Android phone be used as a streaming stick? They can, and that’s exactly what I did with an old Samsung Galaxy phone.

Categories: IT General, Technology

Introducing CARE-X: Towards Clinically Useful Radiology VLMs with Auxiliary Supervision, Reward-Aligned Learning, and Tool-Augmented Measurement

Microsoft Research - 3 hours 46 min ago

Research Note: CARE-X is a research model and not a Microsoft product offering or medical device. It has not been cleared or approved by any regulatory authority and is not intended for clinical diagnosis, screening, or patient care. The results described below are retrospective research findings and do not establish the safety, effectiveness, or suitability of CARE-X for any clinical use. References to potential workflows describe areas for future research, not currently available capabilities or recommended uses. 

At a glance
  • The challenge: Chest X-ray interpretation spans diverse tasks that require both expressive report generation and calibrated diagnostic predictions.
  • CARE-X is a unified chest X-ray VLM for diverse clinical interpretation tasks. It combines generation and structured prediction to provide both free-text reasoning and deterministic outputs.
  • CARE-X uses reinforcement learning (DAPO) to reward clinical correctness in a multi-task setting.
  • In a separate research experiment from CARE-X, we paired Qwen3-VL-4B-Instruct with deterministic measurement tools to evaluate whether direct computation could improve performance on measurement-dependent conditions compared with visual approximation alone. 
  • Validated on real-world Indian clinical data from Narayana Health, including rare ICU pathologies and CT-confirmed enlargement conditions.
What radiologists need: Task diversity, flexibility, and clinical fidelity

A clinically useful radiology AI system must support a wide range of tasks, adapt to different workflows, and produce outputs that are medically accurate.

Radiologists and other clinicians use chest X-rays for many different purposes. A clinically useful AI system must be able to support that range of tasks. It may be asked to generate detailed findings and concise impressions for a report, answer questions about the presence, absence, or location of a finding, identify medical devices and assess their placement, or pinpoint exactly where an abnormality appears in an image.

These tasks also require different kinds of outputs, from narrative reports to calibrated diagnostic scores. And above all, they require clinical accuracy. A report could ostensibly be perfectly written yet clinically wrong if it misses a finding, reverses a negation, or misidentifies a location. Certain findings could be trivial in one context and vital to identify in another.

CARE-X was developed as a research model to explore how a unified approach can address these diverse demands. The system combines generative and discriminative capabilities, clinically aligned optimization, and tool-based reasoning to support a broader range of radiology workflows while maintaining clinical fidelity.

Spotlight: Event Series

Microsoft Research Forum

Join us for a continuous exchange of ideas about research in the era of general AI. Watch the latest episodes on demand.

Watch on-demand Opens in a new tab Gaps in current radiology vision-language models

Despite the impressive task breadth of recent models, critical gaps remain between what radiologists need and what current systems deliver:

  1. No calibrated confidence for diagnostic decisions. Generative VLMs predict diagnoses as free text, but they typically do not provide calibrated confidence scores. In clinical settings, confidence matters. Clinicians cannot tune sensitivity–specificity trade-offs across clinical contexts—an important requirement for real-world deployment. Discriminative models provide these properties but lack the flexibility of open-ended generation.
  2. Cross-entropy loss does not optimize clinical fidelity. Standard training methods treat all token-level errors similarly, regardless of their clinical consequences. A coordinate mistake may be penalized no more than a harmless wording change. A “yes” can be flipped to a “no” even though the clinical meaning is completely different. Missing a life-threatening finding may carry the same training penalty as omitting a minor observation. As a result, models are not explicitly optimized for what matters most in patient care.
  3. No capability for measurement-dependent findings. Some radiological findings require more than visual recognition. Radiological signs such as cardiomegaly, mediastinal widening etc. depend on precise measurements. For example, a model may correctly recognize whether a chest radiograph was acquired using an AP or PA view. But determining cardiomegaly requires measuring the cardiac and thoracic widths and determining the cardiothoracic ratio. Those quantities should be measured and computed rather than visually approximated while considering variables such as type of view, exposure, rotation of the patient etc. 

Together, these gaps call for more than a fluent generative model. The system must combine broad task coverage, structured predictions, clinically aligned optimization, and quantitative tools where direct measurement is required.

CARE-X: One model, flexible outputs

CARE-X brings these diverse interpretation capabilities into one model, using generative or dual inference according to the needs of each task:

Task typeWhat CARE-X doesInference modeReport generation: FindingsProduces the detailed findings sectionGenerativeReport generation: ImpressionProduces the concise diagnostic impressionGenerativePresence and negation assessmentDetermines whether a pathology is present or absent and handles negationDual: generative + auxiliary headDisease location assessmentIdentifies where an abnormality appearsGenerativeFine-grained multilabel disease classificationCategorizes abnormalities across multiple labelsGenerativeMultilabel tubes and lines classificationIdentifies visible medical devicesGenerativeAbnormal placement detection of tubes and linesDetermines whether a device is positioned incorrectlyDual: generative + auxiliary headAbnormality phrase groundingLocalizes a described pathological findingDual: generative + auxiliary headAnatomical groundingLocalizes 29 anatomical regionsDual: generative + auxiliary headTable 1: CARE-X task coverage and inference modes

Dual inference means that a single forward pass produces both an autoregressive response and a structured auxiliary-head prediction with a confidence score. This provides free-text flexibility alongside threshold-adjustable outputs for tasks where operating-point control matters.

The CARE-X architecture and training approach

CARE-X is built on a SigLIP2-so400M vision encoder and a Phi-4-mini-instruct (3.8B) language model connected through a lightweight adapter. To support both free-text generation and structured clinical predictions, the model augments the shared language backbone with task-specific auxiliary heads for classification and visual grounding. These heads provide calibrated diagnostic predictions and spatial localization signals while sharing representations with the generative language model. Rather than being trained independently, they are co-trained with the language-modeling objective, allowing structured supervision to enrich shared representations and improve generative performance on the same tasks.

Training. CARE-X uses a three-stage supervised fine-tuning pipeline (vision pre-training, adapter/head training, and LoRA adaptation) followed by DAPO-based reinforcement learning. DAPO optimizes task-specific rewards for clinical reporting, diagnostic accuracy, and spatial grounding quality.

Figure 1. The CARE-X model. (Left) Supervised fine-tuning with task-specific heads — classification, grounding, and language modeling — sharing the same Phi-4-mini-instruct backbone. The classification head outputs calibrated P(Yes)/P(No) scores; the grounding head outputs bounding box coordinate with confidence; the language modeling head generates free-text responses. (Right) DAPO with task-specific rewards for multi-task reinforcement alignment across report generation, grounding, and VQA.  Auxiliary supervision: Structured prediction strengthens generation

A central finding of this work is that co-training discriminative auxiliary heads with a generative VLM enriches shared representations, leading to stronger generative performance on the same tasks while also providing calibrated structured predictions.

Grounding improvements

The auxiliary grounding head consistently improves localization over generative decoding. On anatomical grounding (Chest ImaGenome), mAP and mIoU increase by +28.2 pp and +6.2 pp, while the largest gains occur on phrase grounding (PadChest), with +24.6 pp mAP and +14.1 pp mIoU. The composite spatial loss enhances geometric precision in shared representations.

DAPO bridges the gap to dedicated detection heads

DAPO-trained generative output approaches or exceeds the SFT auxiliary detection head. On Anatomy grounding, CARE-X generative (0.868 mAP) surpasses the SFT detection head (0.865). This is practically significant—it demonstrates that reward-aligned learning can bring autoregressive spatial decoding to parity with structured prediction, offering clinicians a single generative inference mode without requiring auxiliary heads at test time.

Calibrated classification with tunable operating points

Beyond representation enrichment, the classification head offers a distinct deployment advantage: calibrated probability scores with tunable thresholds allow clinicians to shift between high-sensitivity screening and high-specificity confirmation from a single forward pass—a capability purely generative architectures cannot provide.

ModelInference SettingSensitivity ↑PPV ↑F1 ↑CARE-XGenerative0.9320.8950.913CARE-X (Th=0.5)Auxiliary Head0.9430.8850.913CARE-X (Th=0.6)Auxiliary Head0.8550.9270.890CheXOneGenerative0.8780.8540.866MedGemmaGenerative0.7980.8860.839Table 2: Abnormality classification performance on Chest ImaGenome. Adjustable thresholds enable operating-point selection. Strong report generation across four benchmarks

Within the paper’s comparison set, CARE-X achieves the strongest performance on most reported metrics across MIMIC-CXR, IU-Xray, CheXpert-Plus, and ReXGradient. CRIMSON, a held-out metric that evaluates abnormal findings and weights errors by clinical severity, suggests these gains reflect clinically meaningful improvements rather than reward-specific optimization.

Figure 2. CRIMSON scores (↑) for CARE-X against baseline report-generation models across four chest X-ray datasets — ReXGradient, MIMIC-CXR, IU-Xray, and CheXpert-Plus. CARE-X (highlighted) achieves the highest CRIMSON score on every dataset. CARE-X reaches 94% accuracy on ReXVQA

CARE-X ranks first on the ReXrank RexVQA leaderboard (opens in new tab) as of August 2026. On the ReXVQA benchmark (41,007 question–answer pairs across five clinically relevant categories), CARE-X reaches 94% overall accuracy, six percentage points above the next-best publicly reported model. 

Figure 3: ReXVQA accuracy across five findings-quality dimensions — negation, presence, location, differential diagnosis, geometric information, and overall. CARE-X consistently outperforms CheXOne-R1 and MedGemma on every axis, with the largest margins in differential diagnosis, location assessment and negation. Tool-augmented measurement: Interleaving perception and computation

Some radiological findings depend on quantitative measurements rather than visual patterns. In a separate research experiment from CARE-X, we built an inference-time pipeline that combines Qwen3-VL-4B-Instruct with deterministic measurement tools, allowing the model to alternate between image understanding and precise computation. Qwen3-VL-4B-Instruct retains visual access to the radiograph throughout inference, invoking tools to identify anatomical landmarks, compute measurements, and evaluate diagnostic thresholds as needed. This creates a multi-turn reasoning loop that interleaves perception and measurement, enabling the model to combine visual context with exact quantitative evidence before reaching a diagnosis.

Figure 4. Tool-augmented quantitative reasoning pipeline. The orchestrator mediates a multi-turn loop: the VLM reasons over the image (perception), emits structured tool calls, receives deterministic results, and synthesizes the final diagnosis.

Despite requiring no task-specific training, this approach substantially outperforms perception-only inference across all evaluated measurement-based conditions. The results suggest that for threshold-dependent diagnoses, direct computation of clinically defined measurements is more reliable than visual approximation alone.

More broadly, this measurement-augmented approach could augment clinical workflows by expanding the set of quantitative assessments routinely derived from chest radiographs. For example, aortic dilation is not typically quantified on CXR and is often detected only incidentally on CT scans obtained for other indications. As delayed detection can contribute to adverse cardiovascular outcomes, reliable CXR-based screening could enable earlier identification and follow-up of aortic dilation.

ConditionPerception F1Tool F1Δ F1Cardiomegaly74.5696.00+21.4Mediastinal Widening72.6397.47+24.8Aortic Knob Enlargement60.3199.76+39.5Ascending Aorta Enlargement39.33100.00+60.7Descending Aorta Enlargement†28.57100.00+71.4Average+43.6Table 3: Perception-only versus tool-augmented measurement. The average F1 improvement is 43.6 percentage points across five conditions. Validation on Indian clinical data: Rare ICU conditions and CT-confirmed enlargement

Research ethics and data use: The Narayana Health evaluations used de-identified, retrospective clinical data under applicable institutional ethics review and data-use approvals. Narayana Health approved publication of the study results described here. 

Study 1: Inpatient and ICU conditions

To assess real-world generalizability in a research setting, we evaluated CARE-X on 1,047 de-identified chest radiographs from Narayana Health, annotated for five rare, high-acuity conditions with prevalence ranging from 2.6% to 5.2%—reflecting realistic clinical distributions where missed diagnoses carry severe consequences. 

FractureMediastinal ShiftPneumoperitoneumPneumothoraxTubes & Lines Abnormal PlacementModelSens / SpecSens / SpecSens / SpecSens / SpecSens / SpecCheXOne0.41 / 0.900.80 / 0.780.67 / 0.980.85 / 0.720.03 / 0.97MedGemma0.05 / 1.001.00 / 0.530.00 / 1.000.52 / 0.730.18 / 0.87CARE-X0.62 / 0.640.83 / 0.860.89 / 0.940.83 / 0.750.66 / 0.77Table 4: ICU pathology classification on Indian hospital data. CARE-X achieves the most balanced performance.

CARE-X achieves the highest sensitivity in three out of five conditions while maintaining reasonable specificity, demonstrating generalization to low-prevalence clinical settings.

Study 2: CT-confirmed enlargement conditions

In a retrospective study to measure pure recall efficacy, we evaluated measurement-dependent conditions such as mediastinal widening findings including aortic enlargement, hilar mass, and pulmonary artery enlargement on a outpatient cohort of 122 positive cases with CT-confirmed ground truth, avoiding the subjectivity of radiologist consensus on borderline enlargement findings on CXR. In the overlay setting, the VLM receives the original radiograph alongside a second image with condition-relevant anatomical segmentation masks — offering spatial guidance without direct access to measurement tools.

The tool-augmented variant reached 94.26% recall, a +10.65 percentage-point gain over the best perception-only baseline. Where CT or echocardiography access is limited, reliable triage from a widely available modality like chest X-ray can cut both unnecessary referrals and missed diagnoses.

Figure 5: Recall on the CT-confirmed enlargement cohort across perception-only, overlay-assisted, and tool-augmented inference. (Study 2)

In a related study (accepted at EACTS conference 2026), for mild aortic dilation, the measurement-driven reasoning approach detected 40 of 43 CT-confirmed cases (93% sensitivity), compared to just 5 of 43 (12%) identified on the initial radiology reads, where aortic enlargement is usually not the primary indication for the chest X-ray. This corresponds to 35 additional mild cases that were surfaced but missed during the initial CXR interpretation. These results suggest that explicit quantitative measurements may help identify borderline enlargement that is difficult to assess through visual inspection alone. 

What this does and doesn’t show 

These numbers are all recall, i.e., how many true positives we catch. This was the focus of the initial study because, in triage, a missed diagnosis is typically the costlier failure mode, and CT-confirmed ground truth gave us a clean way to measure it without relying on radiologist consensus for the difficult cases. 

Recall, however, captures only one dimension of diagnostic performance. A model that flags everything achieves perfect recall and is useless in practice. An extended study is underway that includes CT-confirmed negative cohorts as well. Preliminary results are promising, and further studies are planned to explicitly evaluate the viability of quantitative aortic measurements on chest X-ray as a screening tool for aortic dilation. 

CARE-X: Toward clinically useful radiology AI

CARE-X demonstrates that discriminative and generative objectives can be effectively combined within a unified radiology AI model. By jointly training classification, grounding, and language capabilities, the model supports both flexible report generation and calibrated, threshold-adjustable predictions. The separate measurement study further highlights a practical division of labor between learned reasoning and deterministic computation: the VLM provides visual understanding and identifies relevant evidence, while measurement-dependent diagnoses are computed through transparent, tool-based calculations. Retrospective evaluation on clinically challenging Narayana Health cohorts provides encouraging evidence of the potential of this approach for real-world radiology applications. The clinical relevance of this research is underscored by the selection of the AI-based aortic dilatation screening application as a finalist for showcase at the IHF Innovation Hub, World Hospital Congress 2026, recognizing its potential to support earlier detection and clinical decision-making in cardiovascular care. 

Looking ahead, CARE-X can be extended beyond its current capabilities through structured report generation, richer differential diagnosis support, and tighter integration of tools within the model itself. The framework could also benefit from incorporating broader clinical context, including laboratory results and patient history, enabling more comprehensive clinical reasoning. 

CARE-X is a research model, not a Microsoft product offering or medical device. It has not been cleared or approved by any regulatory authority and is not intended or validated for clinical diagnosis, screening, patient care, or clinical decision-making. The results described are retrospective research findings and do not establish safety, effectiveness, or suitability for clinical use.  

Paper co-authors:

Mercy Ranjit, Anirban Porya (opens in new tab), Niharika Vadlamudi (opens in new tab), Nikhilesh E (opens in new tab), Sathvik Joel (opens in new tab), Prasanth V V (opens in new tab), Tanuja Ganu, Abhyuday Swamy (opens in new tab), Pranay Umredkar (opens in new tab), Pradeep Narayan (opens in new tab), Vivek Rajagopal (opens in new tab)

Collaborators: Medha AI (opens in new tab), Narayana Health (opens in new tab)

Opens in a new tab

The post Introducing CARE-X: Towards Clinically Useful Radiology VLMs with Auxiliary Supervision, Reward-Aligned Learning, and Tool-Augmented Measurement appeared first on Microsoft Research.

Categories: Microsoft

Introducing CARE-X: Towards Clinically Useful Radiology VLMs with Auxiliary Supervision, Reward-Aligned Learning, and Tool-Augmented Measurement

Microsoft Research - 3 hours 46 min ago

Research Note: CARE-X is a research model and not a Microsoft product offering or medical device. It has not been cleared or approved by any regulatory authority and is not intended for clinical diagnosis, screening, or patient care. The results described below are retrospective research findings and do not establish the safety, effectiveness, or suitability of CARE-X for any clinical use. References to potential workflows describe areas for future research, not currently available capabilities or recommended uses. 

At a glance
  • The challenge: Chest X-ray interpretation spans diverse tasks that require both expressive report generation and calibrated diagnostic predictions.
  • CARE-X is a unified chest X-ray VLM for diverse clinical interpretation tasks. It combines generation and structured prediction to provide both free-text reasoning and deterministic outputs.
  • CARE-X uses reinforcement learning (DAPO) to reward clinical correctness in a multi-task setting.
  • In a separate research experiment from CARE-X, we paired Qwen3-VL-4B-Instruct with deterministic measurement tools to evaluate whether direct computation could improve performance on measurement-dependent conditions compared with visual approximation alone. 
  • Validated on real-world Indian clinical data from Narayana Health, including rare ICU pathologies and CT-confirmed enlargement conditions.
What radiologists need: Task diversity, flexibility, and clinical fidelity

A clinically useful radiology AI system must support a wide range of tasks, adapt to different workflows, and produce outputs that are medically accurate.

Radiologists and other clinicians use chest X-rays for many different purposes. A clinically useful AI system must be able to support that range of tasks. It may be asked to generate detailed findings and concise impressions for a report, answer questions about the presence, absence, or location of a finding, identify medical devices and assess their placement, or pinpoint exactly where an abnormality appears in an image.

These tasks also require different kinds of outputs, from narrative reports to calibrated diagnostic scores. And above all, they require clinical accuracy. A report could ostensibly be perfectly written yet clinically wrong if it misses a finding, reverses a negation, or misidentifies a location. Certain findings could be trivial in one context and vital to identify in another.

CARE-X was developed as a research model to explore how a unified approach can address these diverse demands. The system combines generative and discriminative capabilities, clinically aligned optimization, and tool-based reasoning to support a broader range of radiology workflows while maintaining clinical fidelity.

Azure AI Foundry Labs

Get a glimpse of potential future directions for AI, with these experimental technologies from Microsoft Research.

Azure AI Foundry Opens in a new tab Gaps in current radiology vision-language models

Despite the impressive task breadth of recent models, critical gaps remain between what radiologists need and what current systems deliver:

  1. No calibrated confidence for diagnostic decisions. Generative VLMs predict diagnoses as free text, but they typically do not provide calibrated confidence scores. In clinical settings, confidence matters. Clinicians cannot tune sensitivity–specificity trade-offs across clinical contexts—an important requirement for real-world deployment. Discriminative models provide these properties but lack the flexibility of open-ended generation.
  2. Cross-entropy loss does not optimize clinical fidelity. Standard training methods treat all token-level errors similarly, regardless of their clinical consequences. A coordinate mistake may be penalized no more than a harmless wording change. A “yes” can be flipped to a “no” even though the clinical meaning is completely different. Missing a life-threatening finding may carry the same training penalty as omitting a minor observation. As a result, models are not explicitly optimized for what matters most in patient care.
  3. No capability for measurement-dependent findings. Some radiological findings require more than visual recognition. Radiological signs such as cardiomegaly, mediastinal widening etc. depend on precise measurements. For example, a model may correctly recognize whether a chest radiograph was acquired using an AP or PA view. But determining cardiomegaly requires measuring the cardiac and thoracic widths and determining the cardiothoracic ratio. Those quantities should be measured and computed rather than visually approximated while considering variables such as type of view, exposure, rotation of the patient etc. 

Together, these gaps call for more than a fluent generative model. The system must combine broad task coverage, structured predictions, clinically aligned optimization, and quantitative tools where direct measurement is required.

CARE-X: One model, flexible outputs

CARE-X brings these diverse interpretation capabilities into one model, using generative or dual inference according to the needs of each task:

Task typeWhat CARE-X doesInference modeReport generation: FindingsProduces the detailed findings sectionGenerativeReport generation: ImpressionProduces the concise diagnostic impressionGenerativePresence and negation assessmentDetermines whether a pathology is present or absent and handles negationDual: generative + auxiliary headDisease location assessmentIdentifies where an abnormality appearsGenerativeFine-grained multilabel disease classificationCategorizes abnormalities across multiple labelsGenerativeMultilabel tubes and lines classificationIdentifies visible medical devicesGenerativeAbnormal placement detection of tubes and linesDetermines whether a device is positioned incorrectlyDual: generative + auxiliary headAbnormality phrase groundingLocalizes a described pathological findingDual: generative + auxiliary headAnatomical groundingLocalizes 29 anatomical regionsDual: generative + auxiliary headTable 1: CARE-X task coverage and inference modes

Dual inference means that a single forward pass produces both an autoregressive response and a structured auxiliary-head prediction with a confidence score. This provides free-text flexibility alongside threshold-adjustable outputs for tasks where operating-point control matters.

The CARE-X architecture and training approach

CARE-X is built on a SigLIP2-so400M vision encoder and a Phi-4-mini-instruct (3.8B) language model connected through a lightweight adapter. To support both free-text generation and structured clinical predictions, the model augments the shared language backbone with task-specific auxiliary heads for classification and visual grounding. These heads provide calibrated diagnostic predictions and spatial localization signals while sharing representations with the generative language model. Rather than being trained independently, they are co-trained with the language-modeling objective, allowing structured supervision to enrich shared representations and improve generative performance on the same tasks.

Training. CARE-X uses a three-stage supervised fine-tuning pipeline (vision pre-training, adapter/head training, and LoRA adaptation) followed by DAPO-based reinforcement learning. DAPO optimizes task-specific rewards for clinical reporting, diagnostic accuracy, and spatial grounding quality.

Figure 1. The CARE-X model. (Left) Supervised fine-tuning with task-specific heads — classification, grounding, and language modeling — sharing the same Phi-4-mini-instruct backbone. The classification head outputs calibrated P(Yes)/P(No) scores; the grounding head outputs bounding box coordinate with confidence; the language modeling head generates free-text responses. (Right) DAPO with task-specific rewards for multi-task reinforcement alignment across report generation, grounding, and VQA.  Auxiliary supervision: Structured prediction strengthens generation

A central finding of this work is that co-training discriminative auxiliary heads with a generative VLM enriches shared representations, leading to stronger generative performance on the same tasks while also providing calibrated structured predictions.

Grounding improvements

The auxiliary grounding head consistently improves localization over generative decoding. On anatomical grounding (Chest ImaGenome), mAP and mIoU increase by +28.2 pp and +6.2 pp, while the largest gains occur on phrase grounding (PadChest), with +24.6 pp mAP and +14.1 pp mIoU. The composite spatial loss enhances geometric precision in shared representations.

DAPO bridges the gap to dedicated detection heads

DAPO-trained generative output approaches or exceeds the SFT auxiliary detection head. On Anatomy grounding, CARE-X generative (0.868 mAP) surpasses the SFT detection head (0.865). This is practically significant—it demonstrates that reward-aligned learning can bring autoregressive spatial decoding to parity with structured prediction, offering clinicians a single generative inference mode without requiring auxiliary heads at test time.

Calibrated classification with tunable operating points

Beyond representation enrichment, the classification head offers a distinct deployment advantage: calibrated probability scores with tunable thresholds allow clinicians to shift between high-sensitivity screening and high-specificity confirmation from a single forward pass—a capability purely generative architectures cannot provide.

ModelInference SettingSensitivity ↑PPV ↑F1 ↑CARE-XGenerative0.9320.8950.913CARE-X (Th=0.5)Auxiliary Head0.9430.8850.913CARE-X (Th=0.6)Auxiliary Head0.8550.9270.890CheXOneGenerative0.8780.8540.866MedGemmaGenerative0.7980.8860.839Table 2: Abnormality classification performance on Chest ImaGenome. Adjustable thresholds enable operating-point selection. Strong report generation across four benchmarks

Within the paper’s comparison set, CARE-X achieves the strongest performance on most reported metrics across MIMIC-CXR, IU-Xray, CheXpert-Plus, and ReXGradient. CRIMSON, a held-out metric that evaluates abnormal findings and weights errors by clinical severity, suggests these gains reflect clinically meaningful improvements rather than reward-specific optimization.

Figure 2. CRIMSON scores (↑) for CARE-X against baseline report-generation models across four chest X-ray datasets — ReXGradient, MIMIC-CXR, IU-Xray, and CheXpert-Plus. CARE-X (highlighted) achieves the highest CRIMSON score on every dataset. CARE-X reaches 94% accuracy on ReXVQA

CARE-X ranks first on the ReXrank RexVQAleaderboard (opens in new tab) as of August 2026. On the ReXVQA benchmark (41,007 question–answer pairs across five clinically relevant categories), CARE-X reaches 94% overall accuracy, six percentage points above the next-best publicly reported model. 

Figure 3: ReXVQA accuracy across five findings-quality dimensions — negation, presence, location, differential diagnosis, geometric information, and overall. CARE-X consistently outperforms CheXOne-R1 and MedGemma on every axis, with the largest margins in differential diagnosis, location assessment and negation. Tool-augmented measurement: Interleaving perception and computation

Some radiological findings depend on quantitative measurements rather than visual patterns. In a separate research experiment from CARE-X, we built an inference-time pipeline that combines Qwen3-VL-4B-Instruct with deterministic measurement tools, allowing the model to alternate between image understanding and precise computation. Qwen3-VL-4B-Instruct retains visual access to the radiograph throughout inference, invoking tools to identify anatomical landmarks, compute measurements, and evaluate diagnostic thresholds as needed. This creates a multi-turn reasoning loop that interleaves perception and measurement, enabling the model to combine visual context with exact quantitative evidence before reaching a diagnosis.

Figure 4. Tool-augmented quantitative reasoning pipeline. The orchestrator mediates a multi-turn loop: the VLM reasons over the image (perception), emits structured tool calls, receives deterministic results, and synthesizes the final diagnosis.

Despite requiring no task-specific training, this approach substantially outperforms perception-only inference across all evaluated measurement-based conditions. The results suggest that for threshold-dependent diagnoses, direct computation of clinically defined measurements is more reliable than visual approximation alone.

More broadly, this measurement-augmented approach could augment clinical workflows by expanding the set of quantitative assessments routinely derived from chest radiographs. For example, aortic dilation is not typically quantified on CXR and is often detected only incidentally on CT scans obtained for other indications. As delayed detection can contribute to adverse cardiovascular outcomes, reliable CXR-based screening could enable earlier identification and follow-up of aortic dilation.

ConditionPerception F1Tool F1Δ F1Cardiomegaly74.5696.00+21.4Mediastinal Widening72.6397.47+24.8Aortic Knob Enlargement60.3199.76+39.5Ascending Aorta Enlargement39.33100.00+60.7Descending Aorta Enlargement†28.57100.00+71.4Average+43.6Table 3: Perception-only versus tool-augmented measurement. The average F1 improvement is 43.6 percentage points across five conditions. Validation on Indian clinical data: Rare ICU conditions and CT-confirmed enlargement

Research ethics and data use: The Narayana Health evaluations used de-identified, retrospective clinical data under applicable institutional ethics review and data-use approvals. Narayana Health approved publication of the study results described here. 

Study 1: Inpatient and ICU conditions

To assess real-world generalizability in a research setting, we evaluated CARE-X on 1,047 de-identified chest radiographs from Narayana Health, annotated for five rare, high-acuity conditions with prevalence ranging from 2.6% to 5.2%—reflecting realistic clinical distributions where missed diagnoses carry severe consequences. 

FractureMediastinal ShiftPneumoperitoneumPneumothoraxTubes & Lines Abnormal PlacementModelSens / SpecSens / SpecSens / SpecSens / SpecSens / SpecCheXOne0.41 / 0.900.80 / 0.780.67 / 0.980.85 / 0.720.03 / 0.97MedGemma0.05 / 1.001.00 / 0.530.00 / 1.000.52 / 0.730.18 / 0.87CARE-X0.62 / 0.640.83 / 0.860.89 / 0.940.83 / 0.750.66 / 0.77Table 4: ICU pathology classification on Indian hospital data. CARE-X achieves the most balanced performance.

CARE-X achieves the highest sensitivity in three out of five conditions while maintaining reasonable specificity, demonstrating generalization to low-prevalence clinical settings.

Study 2: CT-confirmed enlargement conditions

In a retrospective study to measure pure recall efficacy, we evaluated measurement-dependent conditions such as mediastinal widening findings including aortic enlargement, hilar mass, and pulmonary artery enlargement on a outpatient cohort of 122 positive cases with CT-confirmed ground truth, avoiding the subjectivity of radiologist consensus on borderline enlargement findings on CXR. In the overlay setting, the VLM receives the original radiograph alongside a second image with condition-relevant anatomical segmentation masks — offering spatial guidance without direct access to measurement tools.

The tool-augmented variant reached 94.26% recall, a +10.65 percentage-point gain over the best perception-only baseline. Where CT or echocardiography access is limited, reliable triage from a widely available modality like chest X-ray can cut both unnecessary referrals and missed diagnoses.

Figure 5: Recall on the CT-confirmed enlargement cohort across perception-only, overlay-assisted, and tool-augmented inference. (Study 2)

In a related study (accepted at EACTS conference 2026), for mild aortic dilation, the measurement-driven reasoning approach detected 40 of 43 CT-confirmed cases (93% sensitivity), compared to just 5 of 43 (12%) identified on the initial radiology reads, where aortic enlargement is usually not the primary indication for the chest X-ray. This corresponds to 35 additional mild cases that were surfaced but missed during the initial CXR interpretation. These results suggest that explicit quantitative measurements may help identify borderline enlargement that is difficult to assess through visual inspection alone. 

What this does and doesn’t show 

These numbers are all recall, i.e., how many true positives we catch. This was the focus of the initial study because, in triage, a missed diagnosis is typically the costlier failure mode, and CT-confirmed ground truth gave us a clean way to measure it without relying on radiologist consensus for the difficult cases. 

Recall, however, captures only one dimension of diagnostic performance. A model that flags everything achieves perfect recall and is useless in practice. An extended study is underway that includes CT-confirmed negative cohorts as well. Preliminary results are promising, and further studies are planned to explicitly evaluate the viability of quantitative aortic measurements on chest X-ray as a screening tool for aortic dilation. 

CARE-X: Toward clinically useful radiology AI

CARE-X demonstrates that discriminative and generative objectives can be effectively combined within a unified radiology AI model. By jointly training classification, grounding, and language capabilities, the model supports both flexible report generation and calibrated, threshold-adjustable predictions. The separate measurement study further highlights a practical division of labor between learned reasoning and deterministic computation: the VLM provides visual understanding and identifies relevant evidence, while measurement-dependent diagnoses are computed through transparent, tool-based calculations. Retrospective evaluation on clinically challenging Narayana Health cohorts provides encouraging evidence of the potential of this approach for real-world radiology applications. The clinical relevance of this research is underscored by the selection of the AI-based aortic dilatation screening application as a finalist for showcase at the IHF Innovation Hub, World Hospital Congress 2026, recognizing its potential to support earlier detection and clinical decision-making in cardiovascular care. 

Looking ahead, CARE-X can be extended beyond its current capabilities through structured report generation, richer differential diagnosis support, and tighter integration of tools within the model itself. The framework could also benefit from incorporating broader clinical context, including laboratory results and patient history, enabling more comprehensive clinical reasoning. 

CARE-X is a research model, not a Microsoft product offering or medical device. It has not been cleared or approved by any regulatory authority and is not intended or validated for clinical diagnosis, screening, patient care, or clinical decision-making. The results described are retrospective research findings and do not establish safety, effectiveness, or suitability for clinical use.  

Paper co-authors:

Mercy Ranjit, Anirban Porya (opens in new tab), Niharika Vadlamudi (opens in new tab), Nikhilesh E (opens in new tab), Sathvik Joel (opens in new tab), Prasanth V V (opens in new tab), Tanuja Ganu, Abhyuday Swamy (opens in new tab), Pranay Umredkar (opens in new tab), Pradeep Narayan (opens in new tab), Vivek Rajagopal (opens in new tab)

Collaborators: Medhai AI, Narayana Health (opens in new tab)

Opens in a new tab

The post Introducing CARE-X: Towards Clinically Useful Radiology VLMs with Auxiliary Supervision, Reward-Aligned Learning, and Tool-Augmented Measurement appeared first on Microsoft Research.

Categories: Microsoft

Penpot vs. Canva: Why I switched to a self-hosted alternative

How-To Geek - 4 hours 16 min ago

Image editors and desktop publishing tools keep getting better all the time. While I love a good dedicated publishing tool like Scribus, sometimes I need something quick to test layouts or make images/videos fast. There are plenty of image editing tools out there, but Canva and PenPot are two of the most powerful that I've used. So I decided to compare them and see how they both fit into my workflow.

Categories: IT General, Technology

Keep a freshly-manicured lawn with $700 off the Ecovacs Goat A3000 robotic lawn mower

Mashable - 4 hours 22 min ago

SAVE $700.99: As of August 11, get the Ecovacs Goat A3000 Robotic Lawn Mower for $1,799 at Amazon, down from its usual price of $2,499.99. That's a discount of 28%.

Opens in a new window Credit: Amazon Ecovacs Goat A3000 robotic lawn mower $2,499.99 at Amazon
  Get Deal

If the idea of having to spend time mowing your lawn sounds like the worst time possible, it might be a good idea to invest in a robot to handle it for you. Robotic lawn mowers are a great way to keep a tidy lawn without actually having to do anything, and you can get one on sale for an excellent price right now.

As of August 11, get the Ecovacs Goat A3000 Robotic Lawn Mower for $1,799 at Amazon, down from its usual price of $2,499.99. That's $700.99 off and a discount of 28%.

SEE ALSO: The best robot lawn mowers in 2025

This robot vacuum requires no perimeter wire installation and lets you go wire-free thanks to the HoloScope 360 Dual-LiDAR navigation system and AI camera. It can handle itself while mapping your yard, up to 3/4 acre, and can remain accurate up to less than an inch even when dealing with less than optimal conditions.

It uses a TruEdge trimmer to keep borders around your lawn clean-looking, so you don't have overgrown grass and weeds around your flower beds, driveway, or other lawn decor. And with a massive 7,500 mAh battery that can be charged back up in just over an hour, it can get the job done without you having to sit and wait for it to juice back up for too long.

Though we're heading into fall, you're still going to have to contend with keeping your lawn fresh and clean. Make it easier on yourself with this robot and write off one more home task that you won't have to deal with.

Categories: IT General, Technology

I turned my Raspberry Pi bird detector into a game that teaches me bird calls

How-To Geek - 4 hours 31 min ago

Some tech projects generate a lot of information, and that information isn't always used as well as it could be. I have a Raspberry Pi project running bird detection software, and it's identified a large number of species. I decided to see if I could take the audio that it had captured and turn it into an educational game.

Categories: IT General, Technology

Channing Tatum and Gemma Chan lead chilling trailer for Sundance hit Josephine

Mashable - 4 hours 33 min ago

Sumerian Pictures' new trailer for Josephine is a tense experience, teasing a chilling mystery around one question: What did Josephine (newcomer Mason Reeves) see happen in the park? But the film, of course, digs even further.

This Sundance Grand Jury and Audience Award-winner from writer and director Beth de Araújo (Soft & Quiet) is based on her own life experiences. Channing Tatum and Gemma Chan lead the film as parents of eight-year-old Josephine, who witnesses a crime in San Francisco's Golden Gate Park. How can her parents help her process it?

Mashable entertainment editor Kristy Puchko called Josephine "easily the most buzzed-about title out of Sundance."

"Not a tear-jerker but a nuanced family drama that's ripe in emotional intelligence and thought-provoking sequences, Josephine is a hard watch and a must-see," she writes in her Sundance round-up.

The cast also includes Philip Ettinger, Syra McCarthy, and Eleanore Pienta.

Josephine hits cinemas Nov. 20.

Categories: IT General, Technology

Claude now watermarks your generated text for instant detection

How-To Geek - 4 hours 43 min ago

You might want to think carefully the next time you copy-paste text from an AI chatbot. Anthropic has confirmed that its Claude AI models now watermark generated text and files, making it easier to gauge the authenticity of material — with major ramifications for slop and privacy issues.

Categories: IT General, Technology
Syndicate content

eXTReMe Tracker