Blogroll

Google says it fixed Pixel battery drain issues—now Samsung Galaxy owners are complaining

How-To Geek - Thu, 07/30/2026 - 21:49

Google and Samsung are both trying to address battery drain issues on their Android phones, and not always successfully. The two have released July 2026 software updates that have alternately fixed and created battery problems.

Categories: IT General, Technology

AI companies are turning old books into training data, Fahrenheit 451-style

Mashable - Thu, 07/30/2026 - 21:48

AI companies have spent years pulling training material from the internet. Now, as more of the web fills with AI-generated writing, some are looking for text in a place largely untouched by chatbots: used-book shelves.

This Tweet is currently unavailable. It might be loading or has been removed. This Tweet is currently unavailable. It might be loading or has been removed.

Until July 28, ISBNdb, a company that operates an online book database, advertised its ability to source as many as 1 million physical books per order for AI developers. The since-deleted page offered books “tailored to your LLM training needs,” according to 404 Media, including older, specialized, rare, and out-of-print titles gathered from used bookstores and other catalogs.

The company also promised confidentiality. Its marketing materials advertised a “strict NDA on every engagement” and said clients’ identities, strategies, and acquisition targets would not be disclosed. That secrecy attracted attention because converting physical books into AI training data can involve destructive scanning: cutting off their bindings, feeding the loose pages through industrial scanners, and discarding or recycling the originals.

By July 28, ISBNdb had removed the sourcing page and its promise of an NDA. The company said the service had been part of “exploring demand” and that it had “chosen to pivot away from that direction," in a news update.

Still, ISBNdb’s brief sales pitch offered a glimpse into a book-buying operation that one major AI developer has already carried out on a much larger scale. Federal court records show that Anthropic purchased and scanned millions of physical books, while used-book sellers in the United States and Europe have recently reported unusual bulk orders for obscure and out-of-print titles.

The reports have opened several practical questions: Why have older books become so valuable to AI developers, how widespread is destructive scanning, and what happens to the originals once their pages become training data?

Why AI companies are hunting for old books

Older books offer something that has become surprisingly difficult to guarantee online: writing produced entirely by humans.

This Tweet is currently unavailable. It might be loading or has been removed.

Large language models are trained on enormous quantities of text collected from websites, articles, books, code repositories, and other digital sources. Since generative AI tools became widely available, however, the internet has filled with machine-written summaries, product listings, social posts, and articles.

That creates a problem for developers assembling new training datasets. If a model is trained too heavily on material produced by earlier models, it can begin reinforcing their mistakes while losing some of the variety and less common information contained in the original human data. Researchers call the process “model collapse.”

A physical book printed before the current AI boom offers a relatively clean alternative. Unlike a webpage that may have been quietly generated or rewritten by a chatbot, an older book provides a more dependable record of human writing. Books are also edited, structured, and often contain specialized information that cannot easily be found elsewhere online.

ISBNdb leaned heavily on those qualities in its sales pitch.

“Books represent curated, peer-reviewed, domain-specific human knowledge, structured in a way no web crawl can replicate,” the company wrote. “Dense, edited, authoritative.”

ISBNdb also emphasized that many potentially useful titles had never been fully digitized. Its marketing effectively presented the pre-chatbot bookshelf as a reserve of human-created training material that had yet to be mixed with AI output.

That demand did not appear out of nowhere. It follows a much longer effort to turn printed books into searchable digital information.

AI didn’t invent mass book scanning This Tweet is currently unavailable. It might be loading or has been removed.

People have been converting books into digital files for decades. Project Gutenberg began creating electronic versions of public-domain works in 1971 and now offers more than 75,000 free ebooks.

The industrial-scale version arrived with Google Books in 2004. Working with libraries and publishers, Google scanned books and made their contents searchable online. By 2019, the company said it had assembled a collection of more than 40 million books in over 400 languages.

Those scans also helped form HathiTrust, a digital repository created by research libraries in 2008 to preserve their collections and make as much of the material publicly accessible as copyright law allowed.

The project prompted an earlier version of today’s copyright fight. Authors sued Google for scanning copyrighted books without permission, but a federal appeals court ruled in 2015 that the project qualified as fair use. The court found that creating a searchable index and displaying limited snippets gave the books a new purpose without providing readers with a replacement for the originals.

Other digitization programs have emphasized preservation and access. The Internet Archive, for example, says it has digitized more than 25 million books since 2006 using nondestructive scanning designed to keep the bound volumes intact.

This Tweet is currently unavailable. It might be loading or has been removed.

But converting a book into data does not necessarily make it publicly accessible. That distinction concerned Charlie D. Becker, a second-generation bookseller whose family runs Becker’s Books in Houston. His store recently received a single order for 70 obscure titles, including a 1995 guide to metropolitan Denver and manuals explaining how to use WordPerfect in 1991.

After investigating, Becker suspected the buyer was using an algorithm to find underpriced books and relist them on Amazon, rather than acquiring them for AI training. Most of the titles were either unavailable on Amazon or listed there for as much as 20 times his store’s price, and the shipments appeared to be going to Fulfillment by Amazon preparation companies.

This Tweet is currently unavailable. It might be loading or has been removed.

Still, Becker said an obscure book can be especially easy to lose because few people consider it worth preserving. Books that fail to sell through Amazon’s fulfillment system may eventually be liquidated, while some titles have little more than scattered listings across private databases to prove they existed.

“Everyone assumes the internet preserved everything,” Becker wrote in a longer account of the orders. “It didn’t.”

Anthropic has already scanned millions of books

The buyers behind the recent bookstore orders remain unclear. The destructive scanning process, however, is no longer hypothetical.

This Tweet is currently unavailable. It might be loading or has been removed.

In 2024, Anthropic hired Tom Turvey, a former Google executive who had worked on partnerships for the Google Books project. According to a June 2025 federal court ruling, Turvey was tasked with helping the company obtain “all the books in the world” for an internal research library.

Turvey initially contacted publishers about licensing their books, but those conversations did not continue. His team then approached major book distributors and retailers about purchasing print copies in bulk.

Anthropic ultimately spent millions of dollars acquiring millions of physical books, many of them used. Service providers removed the books from their bindings, cut their pages to the appropriate size, and fed them through scanners. The searchable PDF files went into Anthropic’s internal library. The paper originals were discarded.

Engineers could then select groups of those books for inclusion in datasets used to train the large language models behind Claude.

This Tweet is currently unavailable. It might be loading or has been removed.

That process became central to a legal battle over how Anthropic obtained its training material. In June 2025, U.S. District Judge William Alsup ruled that the company’s use of books to train Claude was transformative and qualified as fair use under the specific circumstances of the case.

Alsup also found that converting legally purchased print books into digital files could qualify as fair use because Anthropic destroyed each physical copy and replaced it with one internal digital copy. In other words, the company did not keep both versions.

The ruling did not give AI companies blanket permission to copy any book they could find. Alsup drew a sharp distinction between the print books Anthropic had legally purchased and the more than 7 million pirated books the company had downloaded and stored in a permanent digital library. Purchasing physical copies of some titles later did not erase the original piracy, he found.

That distinction eventually became expensive. On July 20, a federal judge approved Anthropic’s $1.5 billion settlement with authors and publishers, resolving claims involving approximately 482,000 pirated books. Eligible rights holders are expected to receive about $3,000 per title, according to Reuters.

For some authors, the payment does not resolve the larger disagreement. Charles Graeber, one of the case’s original plaintiffs, told NPR that he was proud authors had secured a substantial settlement, but said the case had cost him more than two years of time, travel, and professional opportunities.

Fellow plaintiff Andrea Bartz questioned a system that allows companies to train commercial models on legally purchased books without negotiating separate licenses with the people who wrote them.

“The algorithm is being used to essentially try to put us out of a job,” she told NPR.

Anthropic has maintained that training AI models on books is protected by fair use. The court’s ruling nevertheless helps explain why physical books may be especially attractive to developers: A lawfully purchased copy gives the company a much stronger legal position than a file downloaded from a pirate library.

Destroying the original may be a legally useful distinction. For many readers, it is also the most unsettling part of the story.

Who is buying all these books?

Anthropic’s operation is documented in court records. The source of the more recent bookstore orders is harder to pin down.

This Tweet is currently unavailable. It might be loading or has been removed.

One bookseller specializing in uncommon and low-circulation titles told 404 Media that his weekly sales jumped from roughly 20 books during a good week to several hundred after the orders began arriving in April.

The requested titles did not appear to share a subject, author, genre, or language. They did, however, all have International Standard Book Numbers, or ISBNs, leading the seller to suspect that they had been selected through a book database.

The surge was financially helpful and allowed him to clear inventory that might otherwise have remained unsold. He was less enthusiastic about where the books might be going.

“I don’t like the end-use, and I don’t like that uncommon books are being pulped,” he said.

Booksellers in Europe have reported similarly broad requests. An antiquarian bookseller in the Netherlands received a list of approximately 3,000 English-language books organized by ISBN, ranging from an academic study of Irish folklore to a technical book about laser shock peening.

There is no public confirmation that every unusual bulk order came from an AI company or that every book purchased through these orders was destroyed. There is also no evidence that developers are intentionally hunting for the final surviving copies of rare titles.

The uncertainty itself is part of the concern. A company purchasing from a massive ISBN list could sweep up uncommon or out-of-print editions without first checking how many physical copies remain.

Why the story struck a nerve

Once the reports reached social media, they were accompanied by an unsettling visual: a cutting blade moving inch by inch through a book’s spine. Comparisons to Fahrenheit 451 and the Library of Alexandria followed quickly.

This Tweet is currently unavailable. It might be loading or has been removed. This Tweet is currently unavailable. It might be loading or has been removed. This Tweet is currently unavailable. It might be loading or has been removed. This Tweet is currently unavailable. It might be loading or has been removed.

Book destruction carries a particular weight because it has historically represented more than the loss of paper. The 1933 Nazi book burnings targeted works deemed “un-German,” including books by Jewish, pacifist, and left-wing writers. The act became an enduring symbol of censorship and the suppression of ideas, according to the United States Holocaust Memorial Museum.

That symbolism is central to Fahrenheit 451, Ray Bradbury’s 1953 novel about a society where books are outlawed and burned. Bradbury said his warning extended beyond government censorship to television reducing knowledge to digestible fragments and eroding interest in reading. Users have also invoked the Library of Alexandria, another enduring symbol of lost knowledge, although historians believe it declined gradually through political upheaval, reduced support, neglect, and repeated damage rather than disappearing in one catastrophic fire.

Elon Musk joined the conversation on July 27. He wrote on X that he had asked the SpaceXAI team to preserve rare books in a library and scan them “the hard way,” without removing their spines.

This Tweet is currently unavailable. It might be loading or has been removed.

As the discussion spread, users resurfaced a 2011 Cracked article about libraries, universities, and retailers destroying unwanted books years before the current AI boom.

That history has informed a less alarmed response. Some users argued that the books shown in warehouses appeared to be ordinary, unwanted inventory rather than irreplaceable artifacts. If a book would otherwise be recycled without being read again, they asked, could scanning it first preserve something that would have been lost?

This Tweet is currently unavailable. It might be loading or has been removed. This Tweet is currently unavailable. It might be loading or has been removed.

AI adds a complication. The text may survive, but inside a private company’s research library rather than a public archive. A forgotten manual or travel guide may have little resale value, yet still contain exactly what an AI developer wants: edited human writing created before the flood of chatbot output.

That has led authors and publishers to argue that they should have more control over how their work is used, especially when it is helping companies build commercial products. AI developers, meanwhile, continue to argue that training a model is a transformative use of the material rather than a replacement for the original books.

ISBNdb has removed its sourcing page, but the demand behind it remains. The web’s AI problem has sent developers back to the bookshelf. The next chapter will depend on whether they can extract what they need without leaving those shelves any emptier.

Categories: IT General, Technology

I asked Gemini to automate an Excel reporting task—now a single click does all the work

How-To Geek - Thu, 07/30/2026 - 21:30

I've used Gemini to answer Excel questions for a while, but I wanted to see whether it could handle a real workflow. When I asked it to automate a tedious reporting task, I ended up with a reusable VBA tool that could run in any XLSX workbook—but only after testing, refining, and improving the result.

Categories: IT General, Technology

This AI assistant remembers your anniversary so you dont have to

Mashable - Thu, 07/30/2026 - 21:20

AI assistants have been pitched as a solution to overflowing inboxes, messy calendars, and trips that still need booking. But what happens when the task being outsourced is remembering your anniversary?

The startup Orchid introduced its personal assistant with a now-viral cinematic video, posted to X July 28, following a woman whose partner has forgotten their anniversary. After she tells Orchid what happened, the assistant contacts him, books a dinner reservation, arranges a flower delivery, and keeps her updated on his progress.

"Introducing Orchid, the first assistant that actually gets you," the company wrote alongside the video. "It remembers what you forget. Down to the flower."

This Tweet is currently unavailable. It might be loading or has been removed.

By the end, Orchid has transformed the mistake into what looks like a thoughtful anniversary celebration. Viewers were less charmed.

The video received more than 20 million views within its first several days, but it also attracted hundreds of replies from people disturbed by the future it seemed to present.

This Tweet is currently unavailable. It might be loading or has been removed. This Tweet is currently unavailable. It might be loading or has been removed. This Tweet is currently unavailable. It might be loading or has been removed. This Tweet is currently unavailable. It might be loading or has been removed.

The reservation and flower delivery were hardly the most controversial parts. AI tools already help users search for restaurants, make plans, and select gifts; apparently, romance is now just another workflow waiting to be automated.

So, what does Orchid actually do?

Orchid describes itself as a personal assistant that users communicate with through text messages. The company says it can manage email, schedule meetings, prepare users for calls, plan trips, source gifts, handle invoices, and remember personal details.

Users can also ask Orchid to create recurring "habits," including delivering a morning news briefing, checking the weather, tracking meals, sending medication reminders, and automatically checking them into flights.

This Tweet is currently unavailable. It might be loading or has been removed.

The product's primary audience appears to be professional. According to Orchid's FAQ, it was designed for "founders, investors, lawyers, agents, and agency owners" whose time would be better spent on "real work" than administrative tasks. The company's only published case study focuses on helping an executive prepare for client calls.

But the video shows the assistant doing more than organizing someone's calendar. Orchid appears to connect to the users' iMessage accounts in order to communicate separately with both members of the couple, passing information between them and reporting one person's actions to the other.

Orchid co-founder Nizar Abi Zaher told Gizmodo that "Orchid allows multiple people to communicate with the same agent and share context." For now, he said, that feature is "just group chats and also coordinating meetings." In the advertisement, however, the couple appears to be communicating with Orchid through two private conversations.

Although Orchid's homepage promises "a hundred tools, one assistant," its FAQ says the product presently works with Gmail and Google Calendar. Integrations with Outlook, Slack, Notion, and Linear are still on the way.

The company also says the AI acts "only within the limits that you set." Email responses remain in drafts until the user sends them, while calendar holds require approval. It also says it does not sell users' data or use it to train AI models.

A shared assistant raises additional questions. The ad doesn't explain which information each partner has agreed to share, whether either person can access the other's messages, or how Orchid determines what it can report back. At least, not yet. Mashable has reached out to Orchid for a comment.

Social media isn't sold on Orchid's vision

The possibility of using AI to shape a relationship is not especially new. In a 2025 survey of over 5,000 U.S. singles, 26 percent said they were using AI to enhance their dating lives, a 333 percent increase from the previous year. Nearly half of Gen Z singles had used it for tasks such as filtering matches, writing messages, and evaluating their dating habits.

AI has also made its way into established relationships. A 2026 survey of 1,000 married U.S. adults found that 44 percent had used an AI tool for relationship advice, rising to nearly 65 percent among millennials. That growing use has not necessarily made people comfortable with AI's expanding role — about half of U.S. adults believe AI will worsen people's ability to form meaningful relationships.

That discomfort may help explain why Orchid's ad went over so poorly online: "an AI for perpetuating your dead-end relationship that both parties secretly hate," one X user wrote. Another described Orchid as an "assistant for people who want to subcontract all the basic acts of care that make up a relationship." Some users described the concept as "lame," and an advertisement for "adult babies uninterested in the world they live in." One viewer offered a more concise thought: "Tech people need to stop."

This Tweet is currently unavailable. It might be loading or has been removed. This Tweet is currently unavailable. It might be loading or has been removed. This Tweet is currently unavailable. It might be loading or has been removed. This Tweet is currently unavailable. It might be loading or has been removed.

Some people suspect the company understood exactly how the scenario would be received. "I think they knew what they were doing and hoped the ragebait worked," one user wrote. "Unfortunately it did."

The company has said it intentionally built the advertisement around a story instead of creating a traditional product demonstration. "We wanted to build a story around a real situation," the Orchid team explained in a behind-the-scenes video.

This Tweet is currently unavailable. It might be loading or has been removed.

At least one member of Orchid's team appeared to be taking notes. "If you have feedback on Orchid or ideas on what would make you actually use it, my DMs are open!" Lucas Valbuena, an intern at the startup, posted on X as the backlash spread. Meanwhile, other Orchid employees, AI companies, and engineers publicly rallied behind the company.

This Tweet is currently unavailable. It might be loading or has been removed. This Tweet is currently unavailable. It might be loading or has been removed. This Tweet is currently unavailable. It might be loading or has been removed. This Tweet is currently unavailable. It might be loading or has been removed.

The ad also made its way to r/ExecutiveAssistants, where the person who started the discussion said their "heart kind of broke" after learning about Orchid. Along with worrying about the loss of human connection, they asked whether AI could eventually replace people in their field.

The assistants who responded were mostly skeptical that Orchid could replace them. One predicted that products like it would become tools executive assistants use, leaving humans to catch hallucinations and supply the context AI misses.

"Execs are way too lazy to work these out themselves," the commenter added. Another dismissed Orchid as "marketing fluff aimed at the people who could never really afford us," arguing that a reminder app cannot reproduce the emotional intelligence experienced assistants bring to the job.

That skepticism helps explain why Orchid's ad bothered so many. An assistant can book a table and order flowers, but those gestures usually matter because a partner remembered the occasion and chose to make them in the first place. In the video, Orchid does that work while the human receives the credit.

Orchid promises to remember an anniversary "down to the flower." The backlash suggests people still care who remembers it in the first place.

Categories: IT General, Technology

Toyota's off-road-ready SUV has more than doubled its sales this year

How-To Geek - Thu, 07/30/2026 - 21:00

Few vehicles have managed to buck the trend in 2026, with rising prices and shifting buyer preferences making growth increasingly difficult across the industry. Yet one rugged SUV has done exactly that, posting one of the most impressive sales surges of any mainstream model this year.

Categories: IT General, Technology

We found the 6 best laptops for college students going back to school

Mashable - Thu, 07/30/2026 - 20:49

Here at Mashable, we're constantly testing the best laptops based on an exhaustive in-house methodology that combines real-world use with performance benchmarking. In the past two years alone, we've tried over 80 different models across a range of price points.

To determine which of them fit the bill for college students, specifically, I researched the top 10 public universities' hardware recommendations for the upcoming fall 2026 semester. These include processor types, operating system support, RAM and storage minimums, and other spec requirements for different majors. I honed my guidance based on the common threads I noticed.

I also gave special preference to laptops that are long-lasting and portable enough to be toted around campus, and those that are competitively priced for the current market. Laptops aren't cheap right now, but a solid splurge can last you well past graduation. (Take advantage of student discounts whenever possible.)

Based on this analysis, I eventually settled on six top picks that make the best laptops for college students. Whether you're a humanities student, a STEM major, or still undeclared, I'm confident that at least one of my options will be your ideal machine. You can read more about these picks and my research below.

What to look for in a college laptop, based on my research A Windows laptop with a mid-tier Intel Core or AMD Ryzen 7 processor is suitable for most college students. Credit: Haley Henschel / Mashable

Six of this year's 10 top-ranking public universities tell their incoming students to purchase Windows 11 laptops or MacBooks that are less than two years old. I wouldn't go any older than that for the sake of future-proofing.

Most of those schools tell their students to purchase laptops with 16GB of RAM and 512GB of SSD storage at minimum. (Liberal arts majors might be able to get away with 256GB of storage, but you'll probably have to supplement that with an external hard drive.) More RAM and storage is better if your budget allows; more is mandatory if you're an engineering, design, or computer science major. Students in those fields are generally advised to get a laptop with at least 24GB to 32GB of RAM and 1TB of storage.

SEE ALSO: Laptop specs explained: A jargon-free guide to what's inside your computer

You can approach your processor options in a similar way. For Windows laptops, most schools suggest a mid-range CPU like an Intel Core/AMD Ryzen 5 at minimum, and a mid- to high-end Intel Core/AMD Ryzen 7 or 9 chip for more demanding workloads. Several schools recommend tacking on a dedicated GPU for such coursework, too (i.e, Nvidia GeForce RTX/Radeon RX graphics). On the Apple side, the MacBook Air and Pro with the base M5 chip are go-to recs for most students.

Two schools tell their students not to buy ARM-based Windows laptops (with Qualcomm's Snapdragon processors) because they can't run certain software natively, and because they don't support older peripherals like university printers and scanners. One software example is AutoCAD, a popular 2D and 3D design app that engineering and design students rely on heavily. I love a lot of ARM laptops because they're fast and long-lasting, and I will say that their compatibility is improving every year. Still, I've opted to keep them off my list of 2026 picks out of an abundance of caution; stick with Intel and AMD CPUs for now.

Be sure to budget for an extended warranty with accidental damage protection, as suggested by over half of the top 10 public universities. For reference, AppleCare+ for Mac costs $67.99 to $139.99 a year for students, depending on the MacBook model.

What type of laptop should college students buy? Credit: Joe Maldonado / Mashable

This is a question that only your college can answer for sure, as laptop type recommendations can vary by major. For example, the University of Virginia and the University of California, Davis approve MacBooks for their general student populations, but their engineering departments tell certain tracks to avoid them. Likewise, UCLA's Anderson School of Management hardware requirement page says, "Mac computers are acceptable to use as primary computers for study at Anderson. However, please note that some elective course software is only available for Windows. Students are responsible for ensuring compatibility and configuring their Macs accordingly."

I can say for sure that you probably shouldn't buy a Chromebook. Three schools in my research pool discouraged them for some or all majors, and one school — the University of Florida — only recommended them "as supplemental devices." I included a Chromebook in a previous version of this guide as a secondary option for note-taking, but price increases amid the ongoing RAM shortage have made many of them just as expensive as Apple's budget MacBook Neo (if not more so), which is a much nicer and more powerful computer, relatively speaking.

Do you need your own laptop for college?

Yes, you should have your own laptop at college. Most universities let students temporarily borrow laptops through their libraries and/or tech desks. However, these loaners are available on a first-come, first-served basis and wiped upon return (i.e., you can't save anything on them long-term).

I wouldn't rely too heavily on your school's computer lab, either, because you can't take bring those desktops to class or your dorm. Ultimately, owning your own laptop is way more convenient and ensures you'll have the exact specs needed for your major's coursework.

Recent updates to this guide
  • July 30, 2026: I updated this guide with the results of the Dell 14S' battery life test. It was able to loop a video for 34 hours and 29 minutes before dying — very impressive!

  • July 25, 2026: My preferred configuration of the Acer Swift X 14 is back in stock, so I added it to this guide in place of the Acer Swift X 14 AI. They're both smart buys, but the non-AI version is more powerful and lasts much longer.

  • July 18, 2026: I overhauled this guide to the best laptops for college students ahead of the start of the 2026 school year. I added fresh picks from Apple, Acer, and Dell based on all-new testing and research.

Categories: IT General, Technology

One old smart plug is undoing your entire Wi-Fi security upgrade

How-To Geek - Thu, 07/30/2026 - 20:45

WPA3 is one of those router settings that most people will tell you to use. You might have done it and assumed your job was done, but that's not always true. Unfortunately.

Categories: IT General, Technology

Your ISP watches everything you do: Enable this Windows 11 setting to stop it

How-To Geek - Thu, 07/30/2026 - 20:30

Every time you type a website into your browser, your ISP gets a free ticket to watch it. A few Windows 11 privacy settings get talked about a lot, but this one rarely does, even though it's been sitting inside your PC this whole time. It's not a third-party app, and you don't need to install anything new. You only need to know where to look, and honestly, most people never do.

Categories: IT General, Technology

Tech and chip makers lose $1 trillion in massive AI sell-off

Mashable - Thu, 07/30/2026 - 20:27

Investors are thinking twice about the stocks that have most benefited from the AI boom.

The 20 most valuable chip stocks have lost $1.3 trillion over the past week after a big sell-off, according to analysis by CNBC.

According to the outlet's data, Nvidia led in losses after investors liquidated $238 billion since the market closed on Friday. There's perhaps no bigger sign that investors are getting cold feet when it comes to artificial intelligence, as Nvidia has benefited more than any other company from the AI boom. 

Win an Apple Watch in the Mashable Big Guessing Game!

Other companies in the memory space have also taken a big hit. SK Hynix lost $176 billion. Samsung is down $173 billion. Taiwan Semiconductor Manufacturing Co. lost $119 billion. Micron shed $113 billion. And AMD is down $110 billion.

Tech companies have seen demand in memory and storage skyrocket as AI companies buy out supply to power their insatiable compute needs. Due to this, RAM and SSD storage supply has dwindled for everyday consumers. Consumer tech companies like Apple have been forced to institute price hikes on their products as a result.

Despite this, however, investors are seemingly starting to question their AI-related investments.

“This decline appears to be driven largely by sentiment rather than fundamentals,” Morningstar's Chief Equity Strategist Michael Field told CNBC. “Simply put, it’s loss of confidence,” he added. “We continue to see upside in many AI names, but these are growth stocks, and, as such, much of their value comes from cash flows expected far out into the future, which requires a lot of faith from investors.”

Investor concern surrounding AI technology seems rooted in the fact that while these companies make billions of dollars, they're spending way more than they make. According to recent reporting from the Financial Times and Ed Zitron, AI giant OpenAI had a net loss of $38.5 billion last year. Just this week, OpenAI announced that it will spend $750 billion on infrastructure through 2030. Finally, Google recently experienced its first-ever negative cash flow quarter, thanks to spending on AI infrastructure.

Of course, investor sentiment on AI can turn around quickly.

Categories: IT General, Technology

Your Roku can do more than streaming, here are the hidden features most people miss

How-To Geek - Thu, 07/30/2026 - 20:25

I've had Roku devices in my house for years, going all the way back to when it was just a little box sitting next to the TV. At this point, I've got several Roku TVs and a few sticks spread across different rooms, and for the most part, I've always used them the same way. Open an app, pick something to watch, and move on. It works, so I never really thought much about what else was there.

Categories: IT General, Technology

Your old Galaxy phone is just waiting to be converted into a mini PC

How-To Geek - Thu, 07/30/2026 - 20:20

If you own an old Samsung Galaxy phone you aren't using anymore and need a secondary PC, consider converting it into a mini PC instead of buying one. The only prerequisite is support for Samsung DeX. While you might think that a desktop experience on an Android phone would be dragged down by all kinds of limitations and annoyances, I can attest that it's much better than I expected.

Categories: IT General, Technology

Your Windows Snipping Tool has a built-in screen recorder, text extractor, and AI search

How-To Geek - Thu, 07/30/2026 - 20:00

Most people think of the built-in Windows Snipping Tool as the little thing that pops up when you press the "Print Screen" button, but the truth is that this app can do far more than take screenshots.

Categories: IT General, Technology

The used luxury sedan that feels like a six-figure bargain

How-To Geek - Thu, 07/30/2026 - 19:45

The luxury sedan segment has long been dominated by familiar names. Buyers shopping for a flagship model have traditionally gravitated toward the Mercedes-Benz S-Class, BMW 7 Series, Audi A8, or Lexus LS, expecting class-leading comfort, cutting-edge technology, and impeccable craftsmanship.

Categories: IT General, Technology

I turned a $6 ESP32 into a physical "trash day" reminder I can't ignore

How-To Geek - Thu, 07/30/2026 - 19:30

My trash day reminder is one of my favorite automations. Using an integration that pulls the information from an official website, Home Assistant always knows which trash collection is due. Even using spoken reminders, however, we still sometimes forget to put the trash out on the right day, so I decided to build something that would be harder to ignore.

Categories: IT General, Technology

New Android phones aren't worth buying anymore

How-To Geek - Thu, 07/30/2026 - 19:20

Android is in a strange place right now. New phones have never been better, yet it's harder than ever to justify buying the latest one. You're better off keeping your current device until it's no longer up to the job. And when you do need to replace it, buy an older model instead.

Categories: IT General, Technology

Your other streaming apps have 4 great movies Netflix can't touch

How-To Geek - Thu, 07/30/2026 - 19:01

Netflix is arguably the most popular streaming service in the world. However, it's not the only streamer that boasts great movies. HBO Max, Hulu, Paramount+, and Prime Video all have vast libraries with terrific movies in various genres. If you know where to look, you'll find plenty of certified classics.

Categories: IT General, Technology

Echoverse: Deep, evolving environments for computer-use agents

Microsoft Research - Thu, 07/30/2026 - 19:00
Scaling fidelity over sheer count, targeting the capabilities agents actually lack, and evolving with the models they train. At a glance

We built twelve training worlds for computer-use agents: ten deep domain worlds and two capability worlds, each drilling a single control rendered in many forms (date pickers and nested filters). Depth is what makes them worth training on: these worlds reproduce an application’s real behavior, come seeded with realistic data, and keep state coherent across screens and users. Trained on all twelve, a 9B model nearly doubles its base score (36.5% to 67.1%), coming within fourteen points of GPT-5.4. The experiment taught us several lessons: 

A computer-use agent learns the results of what its actions do only where they have real consequences. A click changes saved state, a message reaches a real person, or a page that refuses to move tells the agent its last move did nothing. A screenshot can show what an interface looks like, but only a working world shows what an action caused.

The consequences worth learning from are stateful, and most of them sit behind a login. The work people want automated lives in closed systems: email and chat, banking, health records, the internal consoles for cloud and ML. You cannot train an agent against the live versions of these. Every attempt writes to a real account, there is no reset between tries, and the true state stays hidden behind the screen. So you rebuild the system as a synthetic world where the database is yours: the state is real and changes for real, but it is safe to break, quick to reset, and graded from the data rather than a screenshot.

By a world we mean three things bound together: an environment (the application, its state, and the actions that change it), the tasks that set goals in it, and a verifier that grades the outcome against ground truth. The community is now good at making them: pipelines stand up an application, seed it, generate tasks, and attach verifiers, yielding hundreds of environments and thousands of checkable tasks. This work builds on that progress. However, once worlds are plentiful and its internal structure becomes the bottleneck: regardless of whether state stays coherent across users and screens, workflows keep their dependencies, a weak skill recurs in enough forms to generalize, and success is judged by outcome or by appearance.

Our bet, the one Echoverse tests, is that the real leverage comes less from adding worlds than from a loop that keeps improving the ones you already have. It treats building the environment and training the model as one process, not two stages: run a model in a world, find where it fails, make the world, its tasks, and its verifiers more faithful or more demanding there, train on the sharper signal, and repeat. Ordinary fine-tuning improves only the model. Here the same graded run that measures the model also improves the world that judged it, so a static benchmark saturates while the loop compounds.

Three levers keep that loop productive, none of them raw environment count. Depth: behaviorally faithful worlds for the domains that matter, including the closed and proprietary ones. Capability targeting: narrow worlds built around the exact interaction a model keeps failing. Co-evolution: improving the environment, its tasks, and its verifiers on every graded run, not just the model.

Figure 1: The learning loop: every graded run is read twice. Surviving failures become model training data, and defects in the world, its tasks, or its verifier become repairs. The same graded run that measures the model also sharpens the world. Why synthetic, and why deep?

Open, login-free sites might seem to remove the need for synthetic worlds, but they make a poor training ground for a different reason: they will not hold still. Pages get redesigned, listings and dates roll forward, and hosts throttle or block automated traffic, so a benchmark that is pinned to them drifts, and no two runs face the same site. An occasional evaluation can absorb that; training cannot, since it runs the same task thousands of times and needs the same world each time. A synthetic world is fixed in time and data: the calendar does not move, the seed data does not churn, and a task means the same thing on the thousandth rollout as on the first. We trade a little surface realism for a world we fully control.

Control is only the floor. A world can be perfectly stable and still be hollow, so what earns training time is depth: not its page count but how faithfully it preserves the causal structure of the work.   Five properties set the bar: behavioral fidelity (controls, permissions, and errors follow the product’s logic); coherent state (a sent message appears for its recipient, a cancelled meeting clears both calendars); workflow depth (an early choice constrains what happens later); authoritative verification (application state, not pixels); and domain value (the workflow is worth improving). In the systems that matter most, the difficulty lives in permissions, shared state, and audit histories: exactly the structure a shallow clone skips. Above this bar, more environments add variety; below it, they add noise.

How the Echoverse factory works?

Echoverse is a single pipeline with two outputs: full domain worlds that preserve workflow depth, and capability worlds that vary one diagnosed interaction. Both lean on the fact that we own the database underneath, so success is a property of the app’s own state, not a model’s read of a screenshot.

Building the world

The pipeline expands a handful of seed scenarios into a spec, then compiles it into machine-checkable claims about routes, state, and behavior. Only then does it generate the app: a FastAPI and SQLite backend under a React interface. A fresh app is a hypothesis, not a world: the builder runs every claim against the running environment, repairing the database, backend, or frontend until each passes, then writes a readiness record that separates hard blockers from advisory risks. A world with open blockers does not advance. Depth here is not a promise in a prompt; it is the list of claims the world has been shown to pass.

Growing the corpus

A world that builds cleanly is still not training data. We reground each task on the live database, drawing goals from entities that actually exist, then send every goal through a panel of analyzers: are its entities real, is the goal plausible, does its difficulty match the work, and, the sharpest test, can an agent driving the real UI complete it? That last check runs in the browser, catching goals no interface can satisfy before a model ever sees them. A generated goal is a claim; a solve against the real app is proof.

Every failure becomes an issue tagged by the layer that must change: database, backend, frontend, task text, or verifier. Layer-specific fixers apply the repair, re-check it against the running app, and roll it back if it regresses. The loop re-scores against database ground truth until the pass rate stops climbing, and each surviving task is exported carrying the exact check that grades it. Those tasks become training data through one process: GPT-5.4 solves each task, a verifier keeps the trajectories that pass ground truth, and those become the supervised fine-tuning (SFT) data behind every experiment below.

Building the world and growing the corpus are not two stages but rather one loop: most defects belong to the world, so we re-version the environment with every iteration. Harder tasks expose gaps in the world, and a sturdier world can carry harder tasks, so each round leaves both stronger.

Figure 2: The environment factory: the two loops behind every world. Phase 1 expands a handful of seeds into an app, then repairs the database, backend, and frontend until it passes machine-checkable claims. Phase 2 regrounds tasks on live data, runs a panel of analyzer and layer-specific fixer agents, and re-scores against database ground truth until the pass rate plateaus. Many of those fixes land in the world itself (dashed arrow).  The verifier is grounded in the database

Every task carries its own answer key, a value or a state change minted from the real database by a SQL query at generation, true by construction and re-checked after the agent finishes. A read is graded on semantic equivalence to the stored value ($288 for $287.62 passes); a write on a real before/after database diff, so claiming a ticket was closed fails unless the row flipped; a read_write scores the lower of the two. Grading is hard to game, grounded rather than labelled, and uniform across an EchoStay booking, an EchoForge issue, and an EchoBank transfer.

Full domains carry the workflow

The domains with the most consequential work are the hardest for public benchmarks to reach: closed, proprietary systems where the difficulty lives in permissions, shared state, and history, not layout. A faithful clone has to reproduce that. What matters is not the pixels but that an action’s consequences reach across screens and users, so a task can run a real workflow and be graded on the state it leaves behind.

The ten Echo domains span communication, technical work, regulated records, community, media, and travel. Where a rich public dataset exists we build on it: EchoStay is seeded from InsideAirbnb, so its listings, hosts, reviews, and amenities are real rather than invented, and EchoForum sits on a public forum corpus of 2.55 million comments. Where none exists, as with mail, calendar, banking, and health records, a seeding pipeline generates the state under strict constraints, dense and internally consistent, not a handful of placeholder rows.

Workflow categoryEnvironmentsDepth the world has to carryCommunication & coordinationEchoMail, EchoCalendar, EchoChatShared threads, schedules, participants, permissions, historiesTechnical creation & operationsEchoML, EchoForgeArtifacts, configuration, dependencies, roles, multi-stage changesRegulated records & transactionsEchoBank, EchoCareBalances or records, authorization, audit history, consequential writesCommunity, media & travelEchoForum, EchoTunes, EchoStayPersistent preferences, social state, search, booking, account actionsTable 1: The ten full-domain environments of the Echo family, grouped by the work they represent. Each is a faithful stand-in for a widely used product, named for the workflow rather than the brand.

That accumulated state is what makes an action’s consequences reach across screens and users. A booking in EchoStay moves through search, listing, availability, and payment across roughly 87 routes and 23 tables, but not a single confirmation screen; an EchoMail thread carries intent from draft through delivery, reply, and label state; an EchoCare order writes each change to an audit trail. The tasks are expensive because of it, often five to twenty actions deep, and finished only when the underlying state has changed.

Figure 3: Per-domain detail across the Echo suite. Each ships as a self-contained, fully-interactive clone of the app it models, with its own backend, seeded database, and feature surface. Counts are grounded database state, not mockups. Capability worlds isolate one skill

Not every weakness represents a missing domain; some are caused by a single control that the agent cannot reliably operate. Picture an agent booking a trip: it searches, filters, opens the right listing, then stalls at the date picker, unable to turn “the second week of March” into the right clicks on an unfamiliar calendar. Building another booking site would not fix that. The skill is learned only when the control itself appears in enough forms, and date pickers and nested filter-and-search are ubiquitous on the live web, rendered a hundred different ways, exactly the variability a single deep app cannot supply.

So we isolate the control and widen the interaction, mass-producing it across layouts, states, and constraints, then generating grounded tasks over each. The datepicker world renders one date control as six core widgets across 10 contexts and holds out 10 new unseen ones, from calendar heatmaps to scroll wheels and fiscal-quarter pickers; its hardest tasks turn transcription into reasoning, resolving “the last Thursday of January 2026” or “10 business days after a start date” to one exact, widget-reachable date. The nested-filter world varies 20 widget families and holds out nine compound-panel families as out-of-distribution, grading every submission by whether the filtered results actually meet the requested conditions, judged by the app’s own logic rather than by appearance.

Figure 4: Every widget family the two skills cover, split into training (in-distribution) and evaluation-only (held out): nested filters, 20 families plus 9 held-out compound panels; date pickers, 6 core types across 10 contexts plus 10 held-out widgets. Figure 5: Date pickers and nested filters themed across domains: nested filters over six verticals, from real estate to pet adoption; date pickers over ten contexts, from scheduling to insurance. What deeper, targeted worlds change

More trajectories do not automatically provide more training signal. What matters is depth: whether an episode carries a task through the dependent steps of a real workflow rather than just rehearsing an action in isolation. Two experiments make the difference concrete from opposite ends: one goes deeper on a whole domain, the other narrows to a single broken skill.

Shallow worlds backfire; deep worlds transfer

A shallow world is the cheap option. It stands up fast and looks convincing, but it only rehearses isolated, correct-looking clicks. Train on that and the model will pick up the wrong reflexes, over-stepping and looping and repeating dead actions, because nothing in the easy world ever punished them. A deep world costs more, but its trajectories carry the dependent structure that transfers to the live site.

To isolate that, take two live WebVoyager domains, Allrecipes and Hugging Face, and compare three checkpoints: the base model and two trained on shallow-world and deep-world trajectories built for those domains. The shallow world poses short, self-contained tasks; the deep world poses tasks that run across dependent steps, where an early action changes the state, options, and verification available later. Both give the model the same domain exposure, so only depth differs, and evaluation uses tasks from the public WebVoyager benchmark for these domains, run on the live sites outside any training world.

On Allrecipes, the shallow world pulls the model down, 80.0% to 75.0%; on Hugging Face it stays flat at 48.0%. Only the deep world improves both, lifting Allrecipes to 85.0% and the harder Hugging Face split to 65.0%. With exposure held equal, the gap is depth: the deep model loops less, and of the 37 Hugging Face tasks, those that exhaust their step budget fall from 15 to nine. What separated the two was not how much the model saw, but whether what it saw preserved the structure of the work.

Figure 6: Deep versus shallow worlds for two live WebVoyager domains, with identical domain exposure and different task depth. Deep lifts both; shallow drops below base on Allrecipes and stalls on Hugging Face. Precision about one skill

The datepicker and nested-filter worlds drill exactly the controls our evaluations flagged, and the two skills reinforce each other. Datepicker training lifts datepicker evaluations (in-distribution 60.0% to 82.6%, held-out layouts 34.0% to 54.0%); filter training lifts held-out filters 62.8% to 84.1%. Gains that hold on forms never trained on indicate that the model learned a rule, not a layout. The skills transfer across each other rather than competing: training either one alone still lifts the other, and training both is the best all-rounder on every split. Against GPT-5.4 as a frontier reference, that combined model already edges ahead on nested filters and closes most of the datepicker in-distribution gap, trailing clearly only on held-out datepickers. And the rule reaches the open web, lifting Online-Mind2Web 29.5% to 34.3% on sites it never saw. 

Figure 7: Targeted training, targeted gains: training either date pickers or nested filters lifts both controls, including held-out widgets and compositions neither was trained on, and training both is the best all-rounder on every split. Higher is better. From synthetic worlds to the live web

Three models run through the rest of this section. Base is Qwen3.5-9B given only a handful of synthetic trajectories, just enough to align a general model to the browser action space. Our model is that same 9-billion-parameter network trained on the full synthetic corpus. GPT-5.4 is a far larger frontier model, included as a reference ceiling.

Does the skill survive the open web? We evaluate our model, unchanged, on WebVoyager and Online-Mind2Web, benchmarks it never trained on. They barely overlap with what we built: both are dominated by open, public sites and read-mostly browsing, while our worlds train login-gated, write-heavy workflows. A large jump was never the point; direction is. The frozen model clears base on both, WebVoyager 66.5% to 71.5% and Online-Mind2Web 40.5% to 43.4% (without BrowserBase, 50.9% to 55.6% and 29.5% to 37.2%), reported through BrowserBase because a hosted browser strips the datacenter bot-blocks and rate limits that otherwise depress every agent’s score. With no live-web data in the mix, this is transfer, not memorization.

Figure 8: Synthetic training transfers to the live web. The full-corpus model, on two benchmarks it never trained on, clears base on both; scores run through BrowserBase to remove datacenter bot-blocks.

The modest live-web gain is a coverage effect, not a ceiling: aim at a live domain and it grows. EchoForge, our code-hosting world, is the same kind of app as GitHub, one of the live sites WebVoyager tests. Add EchoForge to the training mix and the live GitHub score climbs 58.5% to 63.4%, with the overall live scores rising too (WebVoyager 50.9% to 52.9%, Online-Mind2Web 29.5% to 31.1%). The average simply reflects that most of what we built sits in domains these benchmarks never touch.

Figure 9: Closing the gap to the frontier, per environment. The green bar is the gain from base to our model; the faded remainder is the distance still to GPT-5.4. Our model surpasses GPT-5.4 on EchoBank and both nested filters and closes most of the gap elsewhere; each model’s exact score is labelled on the right.

The domains we built, most of them closed and login-gated, tell the opposite story. Across all fourteen, the model nearly doubles base, 36.5% to 67.1%, and where base was weakest it climbs three- to nine-fold, with EchoCalendar, EchoML, EchoChat, EchoCare, EchoForge, and EchoForum all moving from single or low double digits into the forties through sixties. That puts a 9-billion-parameter model within fourteen points of GPT-5.4 on the average (67.1% against 80.7%). On EchoMail, EchoBank, and both nested filters, it matches or beats the far larger frontier model outright, trailing by only a few points on in-distribution datepickers. What gets a 9B model this close is not scale but training data that is deep, targeted, and checkable, exactly what the factory is built to produce. 

What scaling buys, and what it doesn’t

We scaled two axes separately: more trajectories through a fixed set of environments, drawn in equal numbers from each, and more distinct environments. They behave differently. More trajectories on the same worlds keep lifting the in-domain average, though the gains keep shrinking, while transfer to the live web flattens outright: from 6,400 to 20,000 trajectories, WebVoyager holds steady (54.8% to 55.6%) and Online-Mind2Web slips (40.1% to 37.2%). Since every point samples the worlds equally, this is no artifact: each environment holds only so much transferable skill, and once a model has drawn it out, more rollouts mostly polish what it already does. 

Scaling environments produces the opposite result. The average keeps climbing as breadth grows, and WebVoyager reaches its best only with the full set. For generalization, the lever is diversity, not volume. A model reaches sites it never saw by training across many kinds of work, not by seeing one kind many more times. 

Even so, scale itself is not the lever on either axis. A large trajectory budget spent on shallow worlds, or graded against the wrong answer, moves the synthetic number and goes nowhere on the live web. What travels is inside each trajectory: depth that preserves a real workflow, targeting that drills the control an agent fails, and database-grounded grading that keeps the signal honest.

Figure 10: Two scaling axes, scored without BrowserBase. Left: more trajectories on a fixed set of worlds, drawn in equal numbers from each, with the x-axis spaced by actual trajectory count. The synthetic average keeps rising, but live-web transfer saturates, WebVoyager flat and Online-Mind2Web slipping past 6,400 trajectories. Right: more environments, where breadth keeps the synthetic average and WebVoyager climbing. Diversity of environments, not trajectory volume, is what carries skill to unseen sites. The model is not the only thing that learns

The score an agent earns is never the model alone. It comes from a coupled stack: the agent, the environment, the task, and the verifier. A zero can mean the agent failed, or the control is broken, or the requested state is impossible, or the verifier checks the wrong thing. Reading every zero as model supervision trains on defects that should have been repaired. So, we read every graded rollout as a test of the whole stack and let the whole stack learn. The environment improves as broken controls and wiring get fixed, the tasks as goals are re-grounded and made harder, the verifier is fixed when it drifts out of sync with the data. Only failures that survive all three become model curriculum.

EchoStay made this visible. Its failures traced to the world, not the agent: a guest-count control silently broke booking tasks, so a correct booking could never register. Fixing it raised the share of those bookings that could be completed at all from 48% to 78%, recovering 15 of the 24 that had been blocked. The same loop finds different faults elsewhere: EchoForum needed frontend fixes and a page-load speedup, which took one failing set of 37 tasks from 0 solved to 36; EchoChat’s verifier had drifted out of sync with the data, and realigning it lifted the share of gradable tasks from 34% to 99%; EchoCare needed one state-wiring fix; EchoForge had the backend logic but no UI control to reach it.

As the world sharpens, the model climbs with it. Re-running the loop on EchoStay across two rounds, the model trained on its corpus more than doubles, from 16.2% to 38.5%, two-thirds of the distance to GPT-5.4’s 50.4%. The model is not the only thing that learns; it is the thing that compounds once everything under it learns.

Figure 11: Co-evolution lifts the model on EchoStay. As the world went from v1 to v2, the model trained on it more than doubled, from 16.2% to 38.5%, a separate measure from the world’s own solve rate. Higher is better. 

That boundary between repairing the world and teaching the model is easy to hold inside a controlled environment, where both are inspectable. The live web erases it: there is no world to repair mid-task, so when an action lands on nothing, correctness rests entirely on the agent noticing and choosing differently. That is the last thing a world has to teach, and where the live web is least forgiving.

From SFT to RL: Turning worlds into RLEs

Every result so far comes from imitation: the 9B model copies the trajectories GPT-5.4 got right. Imitation inherits a ceiling, though: a clean demonstration never shows how to recover from a mistake or when to stop, the failures that break agents in the wild. Reinforcement learning optimizes the outcome we grade and lets the model learn from its own trajectories, not a teacher’s.

But reinforcement learning needs an RL environment (RLE) it can drive at scale. Each rollout needs a reset to a known state, throughput to sample in parallel, and a reward it can trust, and a run replays the same task thousands of times. The live web is not an RLE: it will not reset, so no two rollouts begin alike; it throttles and blocks automated traffic well before RL’s scale; and it exposes no ground truth, only a screenshot a second model must judge, so the reward is as noisy as the judge and a policy learns the judge’s blind spots rather than the task. Echoverse is an RLE by construction. Every world is a self-contained app we snapshot and reset per rollout, run in parallel, and grade from its own database, so the verifier that filtered the SFT data returns a grounded, verifiable reward rather than one inferred from pixels. The same worlds that benchmark an agent train one.

Figure 12: Reinforcement learning on an Echoverse RLE. From the SFT policy we roll out a group of trajectories in one world; each is a sequence of act and execute steps that changes the database. A grader, the same grounded verifier that filtered the SFT data, sits outside the environment and scores each rollout’s final database state into a reward. The group of rewards updates the policy, and the loop repeats across every training world.

We take the SFT model as the starting policy and run RL against five worlds: EchoBank, EchoForge, EchoForum, EchoStay, and EchoTunes. Tasks come from the harder end of each world, where the SFT policy still leaves headroom, and each update draws on several graded rollouts. Each rollout earns two rewards: a trajectory reward from our database-grounded verifier (LLM judge GPT-4.1), and a dense per-step reward from a multimodal judge that grades each screenshot (GPT-4.1 vision). We train on roughly 100 tasks per world beyond the SFT data, for two epochs. On a held-out set of 25 tasks per world, the judged score rises from 58% to 69%. The teacher taught it what to do; the world taught it when to stop, when to recover, and when to give up.

Figure 13: Reinforcement learning on five worlds, over twoepochs. Left: the held-out judge score (25 tasks per world) climbs from 58% to 69%. Right: the critic’s mean score, the RL reward signal, trends up through training. The reward sums a trajectory reward from our database-grounded verifier (LLM judge GPT-4.1) and a dense per-step reward from a multimodal judge (GPT-4.1 vision). Where this leaves us

A world is no longer a fixed benchmark you score against; it is a training surface you keep improving, where the same graded run that measures the model also sharpens the world that judged it. Deep worlds transferred where shallow clones pulled capability down; one widget rebuilt in a hundred forms taught a skill that reached the live web; co-evolution moved both sides at once; and reinforcement against the same worlds pushed the agent past imitation, lifting held-out performance and trimming wasted steps.

The durable advantage is not the largest inventory of synthetic websites. It is a factory that diagnoses what an agent cannot yet do, builds or repairs the world that teaches it, protects the capability already earned, and runs the loop again. The next turns scale three fronts at once. First, more deep worlds for the closed domains public benchmarks cannot reach. Second, more capability worlds for the interactions models keep failing. And, above all, more reinforcement against those grounded worlds: longer runs, harder tasks, and wider reward exploration that push the agent’s behavior and its performance further than imitation ever could. The levers compound: deeper and broader worlds make stronger RL, stronger RL boosts the agent, and every round exposes the next capability to build.

We are releasing a piece of the factory: environment code and graded test tasks for four worlds, two deep domains (EchoStay and EchoForge) and two capability worlds (the datepicker and nested-filter, each with an in-distribution and a held-out split). Every task carries the database-grounded verifier that scores it, so the same worlds can benchmark an agent or train one. Code and tasks: https://aka.ms/echoverse

When worlds grow at the frontier of an agent’s competence, evaluation stops being a scoreboard and becomes the engine that decides what to build next: worlds that keep learning alongside the agents they train.

Acknowledgments

We thank Alexey Taymanov, Andrew Zhao, Aravind Rajeswaran, Corby Rosset, Hussein Mozannar, Luiz Do Valle, Sara Abdali, Spencer Whitehead, Vibhav Vineet, Zach Nussbaum, Yadong Lu, Pashmina Cameron, Rafah Hosn, and Chinmay Karkar for their valuable help, insightful discussions, and continued support throughout this work.

Opens in a new tab

The post Echoverse: Deep, evolving environments for computer-use agents appeared first on Microsoft Research.

Categories: Microsoft

Inside the eBay harassment campaign that led to a $55.7 million settlement

Mashable - Thu, 07/30/2026 - 18:54

Critical coverage usually draws an angry email or two. For Ina and David Steiner, it led to live cockroaches, a fetal pig, and a bloody Halloween mask.

On Monday, July 27, eBay and several former executives agreed to pay $55.7 million to resolve the couple's lawsuit over the 2019 corporate harassment campaign.

The Steiners are the married founders of EcommerceBytes, a news site covering eBay and the broader ecommerce industry. They filed the civil case in 2021 after members of eBay's security team sent them threats, disturbing packages, and unwanted visitors in an effort to influence the site's reporting.

Of the $55.7 million settlement, $48.7 million will go directly to the Steiners. EBay will pay $46.15 million, former CEO Devin Wenig will pay $2 million, former senior vice president Wendy Jones will pay $500,000, and former chief communications officer Steve Wymer will pay $50,000. The remaining $7 million will go to nonprofit organizations. EBay will contribute $6 million, while Wenig will donate another $1 million to a group protecting First Amendment rights in Ina Steiner's name.

The agreement also allows the Steiners to keep talking publicly about the case. It contains no confidentiality provision, a priority for the couple because they wanted the settlement to discourage other corporations from trying to intimidate journalists over critical coverage.

To understand why the Steiners considered that transparency so important, it helps to go back to how the campaign began.

From critical coverage to criminal charges

When eBay's security team began targeting the Steiners in 2019, the couple had already spent two decades covering the company and other online marketplaces, such as Amazon and Etsy. Through EcommerceBytes, they reported on the issues affecting online sellers, including fees, policy changes, and the executives making those decisions.

The internal lead-up to the campaign only became clear years later — the Steiners initially filed their civil lawsuit in July 2021 and amended it in March 2023. Ina Steiner later told Wired in July 2026 that the litigation gave the couple access to roughly 68,000 documents showing how eBay executives discussed the site behind the scenes.

Those records traced a steady escalation. According to the amended complaint, Wenig sent Wymer a link to an April 10, 2019, EcommerceBytes article about his compensation. Wymer responded, "We are going to crush this lady."

The following month, Jones allegedly asked eBay security chief Jim Baugh to address criticism of the company "off the radar" and told him she did not want to know the details. Then, on Aug. 1, 2019, EcommerceBytes published an article questioning Wenig's handling of eBay's litigation against Amazon. Within half an hour, Wenig allegedly told Wymer that if they were ever going to "take her down," referring to Ina Steiner, "now is the time."

This Tweet is currently unavailable. It might be loading or has been removed.

Four days later, the harassment campaign allegedly began. According to eBay's admissions to federal prosecutors, members of its security team targeted the Steiners between Aug. 5 and Aug. 23, 2019. Anonymous accounts on what was then Twitter criticized EcommerceBytes and threatened to show up at the couple's home in Natick, Massachusetts.

The campaign quickly reached their front door. In addition to the cockroaches, fetal pig, and bloody mask, the group allegedly sent live spiders, fly larvae, a funeral wreath, and a book about surviving the death of a spouse. Pornographic magazines addressed to David were allegedly delivered to a neighbor, while Craigslist ads allegedly invited strangers to the Steiners' home for sex, a block party, and an estate sale. One night, an emergency plumber even arrived unannounced.

As the messages and deliveries continued, the language inside eBay remained aggressive. On Aug. 11, Wymer allegedly told Baugh, "I want to see ashes. As long as it takes. Whatever it takes." The group also took the harassment offline — several members of eBay's security team allegedly traveled from California to Massachusetts, followed the Steiners in a rented van, and attempted to install a GPS tracker on their car.

This Tweet is currently unavailable. It might be loading or has been removed.

They also had an unusual plan for how the campaign would end. EBay security manager Brian Gilbert, a former police captain, was allegedly supposed to approach the Steiners and offer to stop the attacks his colleagues were secretly carrying out. EBay would then appear to have solved a problem its own employees had created.

That plan unraveled when the Steiners realized they were being followed and contacted local police. Once members of the group learned they were under investigation, they allegedly made false statements, deleted digital evidence, and falsified records in an attempt to hide eBay's involvement.

This Tweet is currently unavailable. It might be loading or has been removed.

Wenig, Wymer, and Jones were not criminally charged, and Wenig has maintained that he was requesting a communications response and knew nothing about the harassment. EBay's internal investigation found his messages inappropriate but said it uncovered no evidence that he authorized the security team's actions.

Federal prosecutors eventually charged seven former eBay employees and contractors in 2020, all of whom eventually pleaded guilty. Four received prison sentences between July 2021 and October 2022, including Baugh, who was sentenced to 57 months in September 2022. Two others received one year of home confinement later that year. The final defendant, Gilbert, was sentenced in July 2024 to time served and one year of supervised release.

This Tweet is currently unavailable. It might be loading or has been removed.

The consequences also reached eBay itself. In January 2024, federal prosecutors charged the company with six felony offenses, including stalking, witness tampering, and obstruction of justice. EBay admitted to a detailed account of the campaign, paid the maximum penalty of $3 million, and agreed to retain an independent compliance monitor for three years.

The new settlement, though, resolves the Steiners' civil claims against eBay, Wenig, Jones, and Wymer. The couple also reached separate settlements with the other former employees named in the lawsuit, although those terms were not disclosed.

EBay says it has changed

In a public statement published July 28, 2026, the company called what happened “wrong, reprehensible and should never have happened.” It condemned the employees who pleaded guilty and acknowledged the “unprofessional tone” of communications involving Wenig, Wymer, and Jones.

eBay also pointed to the changes it has made since 2019, including bringing in new leaders and strengthening its policies, internal controls, and employee training.

Wenig, meanwhile, continues to maintain that he knew nothing about the campaign. In a statement provided to Mashable, the former CEO said he was “saddened” by what the Steiners endured, “especially because it occurred during my time as CEO of eBay.” He said the harassment “was deliberately done in secret and without my knowledge.”

After the packages, threats, surveillance, criminal cases, and six years of litigation, the campaign still failed at its original goal. EcommerceBytes remains online, and Ina Steiner got to publish the news of eBay's settlement herself. Talk about closure!

Categories: IT General, Technology

What is a super app? Understanding the AI industrys newest trend.

Mashable - Thu, 07/30/2026 - 18:49

Agentic AI. Claws. Vibe coding. AGI.

AI buzzwords come and go pretty quickly these days, and now there's a new AI trend to know: super apps.

What is a super app?

A super app puts a variety of AI tools into one agentic app. This means that AI chatbots (like ChatGPT) exist within the same space as AI agents, i.e. tools that execute multi-step requests with a degree of autonomy.

The idea is simple: Soon, you'll no longer be bouncing between your AI chatbot, coding assistant, and agentic tools. "Super apps" will centralize all of these tools within a single app. A super app, if you will.

Why are people talking about super apps?

In an earning call this week, Microsoft CEO Satya Nadella said that a new super app will be central to Microsoft's upcoming AI strategy.

"This quarter, we also introduced Autopilots, autonomous, long-running agents with full enterprise compliance, including an always-on personal agent powered by OpenClaw," Nadella said. "And this quarter we will bring these Copilot experiences together, including Code, in one 'super app' spanning both consumer and commercial experiences."

Nadella said that the super app would be released by the end of summer 2026.

Microsoft isn't leading the charge on super apps; they're just the latest company to hop on the trend. AI companies, including OpenAI and Anthropic, are also focusing efforts on super apps.

As reported by Reuters, OpenAI has already launched its super app to employees. Meanwhile, competitor Anthropic is also combining its agentic Cowork tool with its Claude chatbot, the company recently announced.

Super apps could also accomplish something else for AI companies. Because users can access apps like InstaCart, Uber, Zillow, and OpenTable directly within ChatGPT, these agentic super apps could absorb as much of your online activity as possible into a single company's app.

Elon Musk reportedly wants to turn X into an everything app as well, and with super apps, his biggest AI rivals appear to be moving in the same direction.

Disclosure: Ziff Davis, Mashable’s parent company, in April 2025 filed a lawsuit against OpenAI, alleging it infringed Ziff Davis copyrights in training and operating its AI systems.

Categories: IT General, Technology

The hybrid crossover that's built to last

How-To Geek - Thu, 07/30/2026 - 18:45

Toyota has spent decades building a reputation for vehicles that last. From dependable powertrains to strong resale values and low running costs, the brand has become one of the most trusted names for buyers who want a vehicle they can keep for years.

Categories: IT General, Technology
Syndicate content

eXTReMe Tracker