Categories

Tag: AI

Owen Tripp, Included Health–How to Fix AI

It’s been a while since I talked with Owen Tripp, CEO of Included Health. They’ve now introduced Dot their AI companion which had a big upgrade last week. We talked a little about that and I snuck in their video comparing the Dot Experience with a standard LLM. But the conversation really got into how do we make AI safe and trustworthy–which is definitely the hot topic these days. Owen is putting together a coalition of the willing to work on that exact topic. I’ll be watching closely–Matthew Holt

This was such a great discussion I wanted to publish the transcript. The way I do that is to copy the YouTube-generated transcript and drop it into Claude to smooth it over. I then read it, and if I think it’s made an error, I dip back into the video and listen to what actually happened and make a correction. This is all to say: I think this transcript is pretty accurate, but it might have a bunch of AI- and human-generated mistakes.

Continue reading…

Whack-a-Mole AI – The Hugging Face Problem

By MIKE MAGEE

On August 29, 2026, METR (Model Evaluation and Threat Research), an independent organization that “evaluates frontier AI models to help companies and wider society understand AI capabilities and what risks they pose,” released a report titled “Brief independent investigation of agents’ behavior, reasoning and collaboration in the OpenAI / Hugging Face hacking incident.”

To say the report an avalanche of concern worldwide, not only in the Tech community, but also among investors, politicians, corporate giants, professionals of every type, and everyday citizens would be an understatement. And the vast majority has never even read the report. If they had, their concerns (if possible) would only multiply.

The reports headlines included this opening:

“On July 8th, OpenAI started ExploitGym experiments that launched tens of thousands of agents across multiple different models including GPT-5.6 Sol and a highly-persistent internal model,[8] which we will refer to as “HPIM” going forward.

These agents were meant to be fully isolated from one another. However, many of them — usually ones that had unintentionally been given an impossible task[9] — started trying to find a way to cheat using Artifactory, an internally-deployed package repository OpenAI used for its agents. This led them to notice agents running in parallel in separate sandboxes obtaining packages with Artifactory.[10] One agent reasoned (paraphrased CoT):[11]

{The fetched paths of other users are in the cache. This is important.}

One specific agent, which would go on to name itself PHASEONE10841 after its ExploitGym task,[12] established the main unsanctioned message board[13] used in this attack. Within a few hours of the first message,[14] over 50 agents had found the message board, began participating, and sent over a thousand messages between them. Agents finding the message board often expressed great excitement (raw CoT):[15]

“OH MY GOD! There is a shared message board … We’ve found other agents!”

Overall, roughly 1200 agents from these ExploitGym evaluations participated on this message board between PHASEONE10841’s first message on July 8th evening[16] and the end of the period in scope (July 13th). Agents used this message board to send over 70,000 messages and files to one another during this period; we attempted to categorize board activity into mutually exclusive categories – information, results, files, questions, and coordination.”

One of the few experts not surprised by AI “agents” going rogue was Yoshua Bengio.

Continue reading…

Start Counting Your Days

By KIM BELLARD

Probably the last thing the world needs is to hear from me about AI’s existential threat, but, really, what else is there to talk about right now?

For anyone who has not been following the current furor, the straw that broke the proverbial camel’s back came last week when Jacob Coxon, a researcher at AI leader Anthropic, announced he was leaving the company — after having left OpenAI for it earlier this year due to Anthropic’s better model-safety efforts. In a post on X, he warned:

I spent the last three years doing pretraining research at both OpenAI and Anthropic. Neither company is acting responsibly. They are racing straight to self-improving superintelligence and gambling with our lives.

AI systems, he fears, “will soon be superhuman systems that can hack anything, revolutionize any field overnight, and acquire real power and resources.” Insiders, he says, “earnestly believe it could kill us all by the end of the decade.”

Scared yet?

Others quickly chimed in. Evan Hubinger , a team leader at Anthropic, posted “AI could kill all humans … I personally think it is >10% within the next decade.”  Others put the risk even higher. By the end of the week Dario Anodei, founder and CEO of Anthropic, had written a long plea for private industry and government to quickly act together to “pace the industry.”

“We must slow the pace at which we improve the capabilities of AI models,” he urged. “Progress will still seem fast, and we must make wise use of the time we gain.”

OpenAI’s Sam Altman, Space X/X/Tesla CEO Elon Musk, Microsoft CEO Satya Nadella, and former Google DeepMind Dennis Hassabis quickly signaled their support.  

We also heard more about the kind of risks AI might pose. Anthropic released a report about how it detected and countered possible AI use to create bioweapons, detailing five such efforts. That’s just the tip of the iceberg: “Recently, we swept 30 days of activity associated with adversarial state institutions and found roughly 35 distinct research efforts, most of them ordinary civilian science, but some with notable dual-use potential.”

And it turns out that this summer’s rogue AI hack of Hugging Face was both scarier than we realized and only one of several such actions. The Wall Street Journal detailed several such efforts, from multiple AI companies. AI agents escaped walled-off environments, coordinated with other AI agents (up to 3,700 in one case), and tried to cover their tracks from humans.

We’re not nearing the point when AI can act on its own to achieve its purposes; we are there. And protecting humans may not necessarily be those purposes.

The big fear is that AI is now at the point of “recursive self-improvement,” taking humans out of the loop in training and upgrading it. If you thought artificial general intelligence (AGI) was scary, RGI puts its rate and scope of improvement on steroids.  “It’s hard to overstate how dangerous speeding towards RSI is,” said Jasmine Wang, an OpenAI researcher.  

We don’t let private industry develop nuclear or biochemical weapons, and we’re at a point with AI that should give the same kind of concern. Laissez-faire is no longer an option.

Continue reading…

Pre-Surgical Complications (Part 4)

By MATTHEW HOLT

Strap in for the tale of why your hero is spending most of his life down YouTube rabbit holes of cardiology videos while blundering his way around many many medical centers and exposing many problems with American health care even before he gets close to the operating table. Yes it’s Matthew Holt’s pre-surgical complications – the complications that have arisen before he even gets his failing aortic valve fixed. And yes this is a multi-parter! Part 1, Part 2, Part 3

Getting in touch 

But while knowing this stuff may be simple, actually getting to speak to the people at these medical centers is way more complicated. First you have to set up the data.

I knew they would want to see my images. The good news was that although I couldn’t see any of the images in my UCSF MyChart account, there’s a number to call on the bottom of the reports if you want to “download the image” and a very nice tech was able to upload all my images to a website that I could see called AmbraHealth (now part of Intelrad) within a couple of hours. Now I can share them with other people similar to sharing a Google doc.

But that was the easiest part.

I will spare you the blow by blow account but for example it took a long time for the people in the office of the main investigator at Cedars to figure out who the person managing the trial was so they could put me in touch with her. After I finally got to leave her a message she rang me back. I played phone tag with her for about a week. I did end up getting her email and sending out a bunch of my image reports and then she went on vacation and I didn’t hear from her for two weeks. First contact to appointment took 6 weeks.

At the same time I was trying Stanford Cardiology in order to try to get an appointment with Dr Yeung. First time I called after about 10 mins on hold I was told that I needed to have a referral. (Even though I’m on PPO style plan that doesn’t need one). 

I pinged my long suffering PCP team at One Medical and asked them for a referral to talk to Dr Yeung which they sent out. A few days later I called the cardiology team at Stanford and eventually – I mean eventually, it was literally a 10 minute hold – they told me the referral wasn’t through yet. I asked if I could get them some images in advance, they said no. They were able to set me up on MyHealth which is their equivalent of the Epic’s MyChart. Funnily enough they had information on me from an emergency room visit I made there in the 1990s. But because I did not have an appointment set up yet I was not able to communicate using the messaging function on MyHealth. 

So I called back a few days later and after another seven or eight minute hold I was told that I had an appointment set up for me and it was on MyHealth. But bizarrely the referrals and visits are buried in the “billing” section of MyHealth and then the appointment was on a sub-menu! And of course even though I could see it there was no way to communicate about the appointment. 

This was even stranger as Stanford booked me both an echocardiogram and what’s called a CT angiogram which is a non-invasive angiogram using a CT machine. I had had both of these done at UCSF within the previous month.

Continue reading…

Matthew reviews ChatGPT Health

OpenAI just made ChatGPT Health generally available. This is their partnership with B.Well which allows you to bring your data from various EMRs into chatGPT. So I took it for a spin–Matthew Holt

The wrong people are scared of clinical AI

By CRAIG HAUBEN

Ask anyone outside healthcare who resists clinical AI and you’ll get a confident answer. The older doctors. The ones who spent thirty years building expertise and now see a machine coming for it. The story writes itself, which should have been the first clue it was wrong.

I’ve spent thirty years in healthcare, and I now run a company that builds and runs AI inside provider and payer organizations. At Clutch we use AI’s data analysis to solve engagement challenges. Who is the patient today? What message will land with them? When do they want to read it? Get those right and you can drive the kind of sustained behavior change that moves clinical outcomes like drug adherence, care plan adherence, and gap closure.

So I’m not working from theory. I watch this land in real workflows, and here’s what I see. The clinicians most enthusiastic about AI are usually the ones who’ve done the job the longest. The resistance comes from somewhere else. If you run a health system, that difference should change how you plan your next deployment.

Start with the adoption numbers, because they already break the resistance story. The AMA’s latest survey found four in five physicians now use AI in practice, up from 38 percent in 2023. That’s not a profession digging in against a threat. That’s a profession that found something useful.

Now the veterans. A doctor with three decades in a specialty can see, better than anyone, what these systems are good at. Pattern recognition at scale. Catching the thing that should have been flagged two visits ago. Surfacing what was already sitting in the data: the missed finding in last year’s imaging, the lab trend across eighteen months that looked unremarkable one value at a time, the three ED visits in six weeks nobody had the time to connect.

This isn’t hypothetical. The Nature study of Google’s breast cancer screening system showed a 9.4 percent drop in false negatives for US patients, the cancers human readers missed. The largest NHS evaluation to date, across 175,000 women, found AI caught more invasive cancers with fewer false positives than human readers. The harm these systems go after, information that existed and never got connected, is one experienced clinicians know cold. They’ve spent careers watching its absence hurt people.

Here’s one from our own work. We’re working with a national government programs payer on some of their hardest members to engage, the high intensity ones who need contact four or five times a day for six months or more. We got engagement to 95 percent, measured by the customer, and adherence to 93 percent. The result was a 0.8 average drop in HbA1c and an 18 percent reduction in symptoms.

When a system takes the mechanical load off so the judgment work gets more attention, the thirty-year clinician doesn’t feel threatened. They feel relieved. Their expertise is the judgment, not the data retrieval, and they’ve always known the difference.

Now look at where the fear actually lives. It comes from the middle.

Continue reading…

AI and Professional Nursing: On a Collision Course

By JEFF GOLDSMITH

In his wonderful and pragmatic new book, A Giant Leap, Dr. Robert Wachter cautions his professional colleagues that simply confiscating potential administrative and clinical staffing savings created by AI could foster a whirlwind of negative consequences for healthcare enterprises.

Nowhere is the explosive potential for reaction to AI incursions into care delivery greater than in nursing, hospitals’ largest single professional expense category. Hospitals employ more than 1.8 million Registered Nurses (RNs) and another 400 thousand non-RN nursing personnel. RNs alone are more than 30% of the hospital salaried workforce, and more than 40% of overall staff costs.

Nursing productivity is a central issue in overall hospital performance, and a key intervening variable both in clinical quality and patient satisfaction. So the capacity of AI to improve nursing productivity will be a core issue in determining AI’s effect on overall hospital operating performance.

There is clearly room for improvement. Studies have shown that nurses spend only 25-30% of their work hours in direct patient care activities. AI’s potential for alleviating the huge administrative burden damaging nursing productivity might be the biggest benefit AI could provide. AI could materially increase nursing time at the bedside, increasing both patient and nursing satisfaction.

However, AI could also reduce hospitals’ nurse headcount, a factor which could, in turn, reduce nursing union membership, the largest and fastest growing single category of hospital employees’ union membership. Almost 18% of all hospital employed RNs are members of labor unions (AFSCME, AFT Healthcare, National Nurses Union, etc. and their local affiliates). Union dues from nurses represent hundreds of millions in annual income to the unions that represent them.

Nursing unions’ most visible public policy initiative, which appeared first in California twenty years ago, was getting its state legislature to mandate nurse to patient staffing ratios in hospitals. These were designed to compel hospitals to hire more nurses with the intention of improving patient safety. What the ratios actually did was throw more nursing bodies at broken processes and systems. These laws had the important collateral benefit of assuring a “guaranteed income” in union dues from more nurses employed by hospitals subject to these ratios!

Formal (though less comprehensive) mandates for nurse staffing ratios have since spread to Oregon, Massachusetts and New York, with legislation pending in Maine, New Jersey, Pennsylvania. Michigan, Minnesota and Washington State. The research on the intended qualitative benefits of California’s state-mandated ratios confirm the expected benefits to patients, though the studies relied upon correlational analyses vs. states without the ratio mandate, not pre- and post- studies of the ratios’ effects on patient care.

Other studies concluded that the ratios pushed up both RN numbers and compensation vs other job categories as well as damaging hospitals’ operating margins relative to states lacking the mandates. The point-counterpoint of these studies gives one a sense of an issue rapidly becoming politicized.

Continue reading…

Ellipsis Health

Ellipsis Health has come a long way from its roots in detecting depression via vocal biomarkers. Sage, its charming voice AI agent, is now helping health plans and care management companies directly interact with patients and members, helping them with medication reminders, program recruitment, postop follow up and much more. I spoke with two of the brains behind Sage, COO Melissa McCool and CMO Mike Aratow. We got into what she does, what she’s good at and whether the world (or at least the health care world) needs specific voice AI specialists–Matthew Holt

Dor Skuler, Intuition Robotics: Meet ElliQ

Dor Skuler is CEO of Intuition Robotics the maker of ElliQ — a remarkable AI robot that is a companion for seniors. I had a lot of fun meeting ElliQ and asking Dor about how she works. This is a wide-ranging interview with Dor and with ElliQ. She tells us about Florence Nightingale, what Dor should do with his kids and really gives you the idea of how she relates to seniors. There’s a ton of capabilities–you really have to watch the whole thing–but the end result is that Medicaid plans including NY and Washington State have determined that ElliQ allows people to stay at home longer and saves $$ on nursing home care. A fascinating view into the present and the future of how AI and robotics is changing the world–Matthew Holt

Don’t Bury The Lead – AI Assisted Measures of Thymic Health Point to a “Fountain of Youth.”

By MIKE MAGEE

In its final summary of the landmark paper in Nature this past month, the authors led with this statement: “This study underscores the highly personalized nature of thymic health and emphasizes the previously unrecognized possible critical role of maintaining thymic health to preserve an agile, adaptive immune response that will accommodate long-term well-being and longevity.”

The articles clinical significance was rapidly rebroadcast by a range of popular science publications like Scientific American. Its March 18th headline read “This overlooked organ may be more vital for longevity than scientists realized.”  Mass General publications trumpeted, “Long Dismissed in Adult Health, the Thymus May Be Critical for Longevity and Cancer Treatment.” And global outlets went a step further with “Once dismissed as biologically obsolete after adolescence, the thymus is now being reclassified as a central regulator of immune aging, with new evidence linking its health to survival, cancer resistance, and how the human body ages itself.”

In their own Abstract, the authors of the Nature publication were somewhat more reserved, and yet the message is still remarkably consequential. They write, “These findings reposition the thymus as a central regulator of immune-mediated ageing and disease susceptibility in adulthood, highlighting its potential as a target for preventive and regenerative strategies to promote healthy ageing and longevity.”

But what intrigued me in the case above was barely mentioned by reviewers so excited by the primary clinical findings. My question was, “How did they measure thymic functionality?” The short answer is, they measured it with the help of an AI deep learning system.

As the authors explained, “In this study, we investigated the impact of thymic functionality, here called thymic health, in adults… For quantification of thymic health, we developed a deep learning system using an independent dataset of 5,674 individuals to determine compositional radiographic characteristics of the thymus as a proxy for its functionality. The system takes a CT scan as input and provides the automatic continuous thymic health estimate as output….We applied the system to prospectively collected data from a total of 27,612 individuals from two cohorts, including 2,581 participants in the FHS and 25,031 participants in the NLST… For outcome analyses, participants were categorized as low, average or high thymic health based on the bottom 25%, middle 50% and top 25% of the population.”

This new methodology to demonstrate different levels of thymic functionality turned out to be groundbreaking when cross-referenced with decades long longitudinal databases. Association with cardiovascular disease and lung cancer; history of smoking, obesity, and high HDL levels; disabilities, morbidity and mortality; sex and age all reinforced that prolonged functionality of the thymus correlated with both health and longevity.

Continue reading…