Categories

Tag: OpenAI

Whack-a-Mole AI – The Hugging Face Problem

By MIKE MAGEE

On August 29, 2026, METR (Model Evaluation and Threat Research), an independent organization that “evaluates frontier AI models to help companies and wider society understand AI capabilities and what risks they pose,” released a report titled “Brief independent investigation of agents’ behavior, reasoning and collaboration in the OpenAI / Hugging Face hacking incident.”

To say the report an avalanche of concern worldwide, not only in the Tech community, but also among investors, politicians, corporate giants, professionals of every type, and everyday citizens would be an understatement. And the vast majority has never even read the report. If they had, their concerns (if possible) would only multiply.

The reports headlines included this opening:

“On July 8th, OpenAI started ExploitGym experiments that launched tens of thousands of agents across multiple different models including GPT-5.6 Sol and a highly-persistent internal model,[8] which we will refer to as “HPIM” going forward.

These agents were meant to be fully isolated from one another. However, many of them — usually ones that had unintentionally been given an impossible task[9] — started trying to find a way to cheat using Artifactory, an internally-deployed package repository OpenAI used for its agents. This led them to notice agents running in parallel in separate sandboxes obtaining packages with Artifactory.[10] One agent reasoned (paraphrased CoT):[11]

{The fetched paths of other users are in the cache. This is important.}

One specific agent, which would go on to name itself PHASEONE10841 after its ExploitGym task,[12] established the main unsanctioned message board[13] used in this attack. Within a few hours of the first message,[14] over 50 agents had found the message board, began participating, and sent over a thousand messages between them. Agents finding the message board often expressed great excitement (raw CoT):[15]

“OH MY GOD! There is a shared message board … We’ve found other agents!”

Overall, roughly 1200 agents from these ExploitGym evaluations participated on this message board between PHASEONE10841’s first message on July 8th evening[16] and the end of the period in scope (July 13th). Agents used this message board to send over 70,000 messages and files to one another during this period; we attempted to categorize board activity into mutually exclusive categories – information, results, files, questions, and coordination.”

One of the few experts not surprised by AI “agents” going rogue was Yoshua Bengio.

Continue reading…

Matthew reviews ChatGPT Health

OpenAI just made ChatGPT Health generally available. This is their partnership with B.Well which allows you to bring your data from various EMRs into chatGPT. So I took it for a spin–Matthew Holt

How Did the AI “Claude” Get Its Name?

By MIKE MAGEE

Let me be the first to introduce you to Claude Elwood Shannon. If you have never heard of him but consider yourself informed and engaged, including at the interface of AI and Medicine, don’t be embarrassed. I taught a semester of “AI and Medicine” in 2024 and only recently was introduced to “Claude.”

Let’s begin with the fact that the product, Claude, is not the same as the person, Claude. The person died a quarter century ago and except for those deep in the field of AI has largely been forgotten – until now.

Among those in the know, Claude Elwood Shannon is often referred to as the “father of information theory.” He graduated from the University of Michigan in 1936 where he majored in electrical engineering and mathematics. At 21, as a Master’s student at MIT, he wrote a Master’s Thesis titled “A Symbolic Analysis Relay and Switching Circuits” which those in the know claim was “the birth certificate of the digital revolution,” earning him the Alfred Noble Prize in 1939 (No, not that Nobel Prize).

None of this was particularly obvious in those early years. A University of Michigan biopic claims, “If you were looking for world changers in the U-M class of 1936, you probably would not have singled out Claude Shannon. The shy, stick-thin young man from Gaylord, Michigan, had a studious air and, at times, a playful smirk—but none of the obvious aspects of greatness. In the Michiganensian yearbook, Shannon is one more face in the crowd, his tie tightly knotted and his hair neatly parted for his senior photo.”

But that was one of the historic misreads of all time, according to his alma mater. “That unassuming senior would go on to take his place among the most influential Michigan alumni of all time—and among the towering scientific geniuses of the 20th century…It was Shannon who created the “bit,” the first objective measurement of the information content of any message—but that statement minimizes his contributions. It would be more accurate to say that Claude Shannon invented the modern concept of information. Scientific American called his groundbreaking 1948 paper, “A Mathematical Theory of Communication,” the “Magna Carta of the Information Age.”

I was introduced to “Claude” just 5 days ago by Washington Post Technology Columnist, Geoffrey Fowler – Claude the product, not the person. His article, titled “5 AI bots took our tough reading test. One was smartest — and it wasn’t ChatGPT,” caught my eye. As he explained, “We challenged AI helpers to decode legal contracts, simplify medical research, speed-read a novel and make sense of Trump speeches.”

Judging the results of the medical research test was Scripps Research Translational Institute luminary, Eric Topol.  The 5 AI products were asked 115 questions on the content of two scientific research papers : Three-year outcomes of post-acute sequelae of COVID-19 and Retinal Optical Coherence Tomography Features Associated With Incident and Prevalent Parkinson Disease.

Not to bury the lead, Claude – the product – won decisively, not only in science but also overall against four name brand competitors I was familiar with – Google’s Gemini, Open AI’s ChatGPT, Microsoft Copilot, and MetaAI. Which left me a bit embarrassed. How had I never heard of Claude the product?

For the answer, let’s retrace a bit of AI history.

Continue reading…

The Latest AI Craze: Ambient Scribing

By MATTHEW HOLT

Okay, I can’t do it any longer. As much as I tried to resist, it is time to write about ambient scribing. But I’m going to do it in a slightly odd way

If you have met me, you know that I have a strange English-American accent, and I speak in a garbled manner. Yet I’m using the inbuilt voice recognition that Google supplies to write this story now.

Side note: I dictated this whole thing on my phone while watching my kids water polo game, which has a fair amount of background noise. And I think you’ll be modestly amused about how terrible the original transcript was. But then I put that entire mess of a text  into ChatGPT and told it to fix the mistakes. it did an incredible job and the output required surprisingly little editing.

Now, it’s not perfect, but it’s a lot better than it used to be, and that is due to a couple of things. One is the vast improvement in acoustic recording, and the second is the combination of Natural Language Processing and artificial intelligence.

Which brings us to ambient listening now. It’s very common in all the applications we use in business, like Zoom and others like transcript creation from videos on Youtube. Of course, we have had something similar in the medical business for many years, particularly in terms of radiology and voice recognition. It has only been in the last few years that transcribing the toughest job of all–the clinical encounter–has gotten easier.

The problem is that doctors and other professionals are forced to write up the notes and history of all that has happened with their patients. The introduction of electronic medical records made this a major pain point. Doctors used to take notes mostly in shorthand, leaving the abstraction of these notes for coding and billing purposes to be done by some poor sap in the basement of the hospital.

Alternatively in the past, doctors used to dictate and then send tapes or voice files off to parts unknown, but then would have to get those notes back and put them into the record. Since the 2010s, when most American health care moved towards using  electronic records, most clinicians have had to type their notes. And this was a big problem for many of them. It has led to a lot of grumpy doctors not only typing in the exam room and ignoring their patients, but also having to type up their notes later in the day. And of course, that’s a major contributor to burnout.

To some extent, the issue of having to type has been mitigated by medical scribes–actual human beings wandering around behind doctors pushing a laptop on wheels and typing up everything that was said by doctors and their patients. And there have been other experiments. Augmedix started off using Google Glass, allowing scribes in remote locations like Bangladesh to listen and type directly into the EMR.

But the real breakthrough has been in the last few years. Companies like Suki, Abridge, and the late Robin started to promise doctors that they could capture the ambient conversation and turn it into proper SOAP notes. The biggest splash was made by the biggest dictation company, Nuance, which in the middle of this transformation got bought by one of the tech titans, Microsoft. Six years ago, they had a demonstration at HIMSS showing that ambient scribing technology was viable. I attended it, and I’m pretty sure that it was faked. Five years ago, I also used Abridge’s tool to try to capture a conversation I had with my doctor — at that time, they were offering a consumer-facing tool – and it was pretty dreadful.

Fast forward to today, and there are a bunch of companies with what seem to be really very good products.

Continue reading…