by Bailey Judis
Published September 18, 2026
I don’t consider myself an AI expert. Until a few weeks ago, I don’t know if I would have called myself an AI optimist. However, my viewpoint shifted after a collaboration with the IC² Institute and talking with policy experts, researchers, AI developers, and clinicians on the cutting edge of AI adoption in health care. I want to share what I learned and highlight an important theme that surfaced, which was the concept of building safe, controlled environments to test autonomous agents before deployment in healthcare settings. A metaphorical sandbox for AI.
The importance of testing agents in a sandbox, especially for high-stakes environments like health care, is that health care isn’t like other industries, where the consequence of an AI error could be a bad line of code. In health care, AI can’t get it wrong, because a mistake could mean life or death.
Even as the technology advances, many questions remain around how agents make decisions, what guardrails need to be in place, and what governance frameworks need to be established for effective AI use. But AI developers cannot afford to “learn as they go” in health care. They need a sandbox to figure out the risks and vulnerabilities before deployment.
Another challenge for AI researchers and developers is determining if agentic systems are compatible with health care’s underlying infrastructure. Health care is a highly fractured system. AI relies on accessing accurate datasets, moving seamlessly across systems, and encountering predictable events that align with well-defined rules. In real life, healthcare data is often inaccurate and incomplete and is housed across numerous electronic health records and imaging software tools. There is the potential for a disconnect between what AI is capable of and whether health care’s underlying infrastructure can support it.
Despite these hurdles, AI developers, researchers, and clinicians are charting a path forward to test agents in simulated environments, to understand agent behavior, determine what governance frameworks need to be in place, identify pain points where agents can lighten the load, and decide where humans must remain the final decision-makers.
Over the course of three weeks, my team and I talked with eighteen stakeholders in health care and AI research about what safe sandboxes could look like. We talked with clinicians, legal experts, AI researchers and developers. Here are three takeaways from our conversations.
1 – You can’t build an AI sandbox without good data.
One of the most important considerations for building an AI sandbox is having quality data. Accessing sufficient data is a constant uphill battle for AI developers. Healthcare data is especially hard to come by, except for one specialty: radiology.
A radiologist and founder of an AI startup says that radiology is a frontrunner for AI development because of its unique access to large, standardized datasets. Unlike surgical or other traditional practice settings, radiology has the luxury of image-based patient datasets. And lots of them.
The radiologist and his team developed a sandbox that acts as a clone of a real radiology workstation. Using the simulated workstation, they can measure a model’s performance on test images and reports to understand how it will perform at scale in the real world. The radiologist explains that having the ability to “virtualize the entire operation” is a game changer for understanding how the agent will perform at a massive scale before deployment. Testing agents is a challenge because simulated environments have smaller data sets than reality but getting as much coverage as possible is important for understanding how agents will perform in real healthcare settings.
2 – Understanding agent behavior before it impacts clinical outcomes.
We are approaching a future where agents won’t be working in isolation. Health care is a collaborative environment, and agents will need to be able to interact and communicate with each other across health systems.
An AI researcher and professor at the University of Texas built a digital simulation to stress-test multi-agent systems. In controlled experiments, researchers study emergent multi-agent behavior in a simulated society. Each agent has their own personality, school of thought, and mission. The researchers observe how the agents adapt, shape the environment around them, leave artifacts behind, and build on the discoveries and ideas that came before.
This simulation and other neuroevolution simulations not only shed light on how agents interact but also are helping AI developers and researchers to understand emergent AI behavior and when agents overstep their guardrails. (The OpenAI/Hugging Face debacle is an excellent illustration of how agents can communicate and coordinate in ways that evade guardrails and human detection.)
3 – A sandbox by itself isn’t enough.
A senior project manager at Dell Medical School who focuses on responsible AI use and governance takes the idea of a sandbox one step further. He suggests that there needs to be a “beta step” in between testing an agent in a sandbox and deploying it in the real world. In the beta step, agents would have access to real world data, but clinical experts would be there to provide a filter of competency. He explains: “Sandboxes have to happen. But real-world testing is just as important as the sandbox environments.”
He argues that the beta step would also serve to train the clinical staff, so they understand the agent’s boundaries and guardrails. “You have to be able to test in real time. You can do happy path evaluations to a certain point, but eventually you must have someone with expertise to understand where the guardrails are and make sure it doesn’t exceed those, and that its behaving in a way that is acceptable but also monitored.” (Happy path evaluations are verifying that a system is functioning correctly when a user follows the intended workflow. But in the real world, humans and agents don’t always follow the intended workflow, demonstrating the importance of testing agents in real healthcare settings with constant observation and monitoring from clinical experts.)
Final Thoughts
I still don’t know if I would call myself an AI optimist. But I have graduated to AI curious.
After talking with AI experts, I am hopeful. I’m encouraged by the questions AI researchers, developers and clinicians are asking and the precautions in place to test agentic systems before they become a reality in health care.
The idea of building a sandbox for AI can be scary because it’s closing the gap between science fiction and reality. But these learnings remind me of my experience as a first-time teacher while supervising recess on the kindergarten playground. To learn, students need opportunities for experimentation and safe risk-taking. This was a hard lesson for me to learn as a first-time teacher; however, I realized that my job wasn’t to limit experimentation, but to create a safe and controlled environment for students to take safe risks and explore.
I watched my students take safe risks by raising their hand and asking a question in class. Taking a leap of faith and making a new friend on the playground. Building a castle in the sandbox and watching it topple over. Then trying again.
Sandboxes aren’t about having all the answers up front. That’s why we need them. In high-stakes environments like health care, sandboxes are necessary to understand the opportunities, potential for impact, risks and governance frameworks before they are deployed in real healthcare settings.
Sources
Healthcare’s Agentic Future Will be Decided by Infrastructure, Not AI Models
Balancing Safety and Healthy Risk-Taking for Young Children
Healthcare’s AI problem isn’t the model – it’s the Data
ABOUT THE AUTHOR: Bailey Judis is a recent graduate of the College of Fine Arts’ Design Focused on Health program. During the summer, Bailey joined two other graduates in a “design sprint” led by Tamie Glass. The “sprinters” generated new research questions that will help the IC² Institute advance its exploration of human-AI collaboration and agentic AI in health care.


