Sign in to view Jonas’ full profile
or
New to LinkedIn? Join now
By clicking Continue to join or sign in, you agree to LinkedIn’s User Agreement, Privacy Policy, and Cookie Policy.
San Francisco, California, United States
Sign in to view Jonas’ full profile
Jonas can introduce you to 10+ people at Handshake
or
New to LinkedIn? Join now
By clicking Continue to join or sign in, you agree to LinkedIn’s User Agreement, Privacy Policy, and Cookie Policy.
8K followers
500+ connections
Sign in to view Jonas’ full profile
or
New to LinkedIn? Join now
By clicking Continue to join or sign in, you agree to LinkedIn’s User Agreement, Privacy Policy, and Cookie Policy.
View mutual connections with Jonas
Jonas can introduce you to 10+ people at Handshake
or
New to LinkedIn? Join now
By clicking Continue to join or sign in, you agree to LinkedIn’s User Agreement, Privacy Policy, and Cookie Policy.
View mutual connections with Jonas
or
New to LinkedIn? Join now
By clicking Continue to join or sign in, you agree to LinkedIn’s User Agreement, Privacy Policy, and Cookie Policy.
Sign in to view Jonas’ full profile
or
New to LinkedIn? Join now
By clicking Continue to join or sign in, you agree to LinkedIn’s User Agreement, Privacy Policy, and Cookie Policy.
About
Welcome back
By clicking Continue to join or sign in, you agree to LinkedIn’s User Agreement, Privacy Policy, and Cookie Policy.
New to LinkedIn? Join now
Activity
8K followers
-
Jonas Mueller shared thisAI promises to cure disease, but can frontier models interpret the images biotech/pharma scientists handle every day? To evaluate this, we built the ATLAS Visual Life Sciences (VIALS) benchmark: 161 visual interpretation tasks from professional life sciences workflows including gel blots, plasmid maps, molecular structures, and phylogenies. These are the images/results that experiments actually produce, not polished figures from papers or textbooks. We benchmarked today’s 10 leading multimodal models. Opus 5, GPT 5.6, and Gemini 3.7 all fail to crack 27% accuracy, and the other models like Kimi K3 & Grok 4.6 perform much worse. In contrast, PhD-level biomedical scientists find these visual interpretation tasks straightforward. AI that cant interpret such images will have limited utility in life sciences R&D, where such artifacts are central to how scientists reason, communicate, and make research decisions. We hope VIALS drives progress toward models that can
-
Jonas Mueller reposted thisJonas Mueller reposted thisOn August 12th, ~100 technical leaders from frontier labs are gathering at a16z's flagship SF venue to discuss one of the most critical questions in AI today. The event "AI’s Next Bottleneck: The State of Training Data" will feature discussions led by: 💬 The Nativity of Today's Data Markets, presented by Engy Ziedan, Chief Scientific Officer and Co-Founder at Protege 💬 Discussion on Bottlenecks in Frontier Applied AI, led by: Michael Bendersky, Research Director at Databricks Jeff Wang, President of New Enterprise a Cognition - Moderated by Shangda Xu, partner at a16z. 💬 Discussion on Bottlenecks in RL Environments, led by: Jonas Mueller, Ph.D., Director of Research at Handshake Jay Ram, CEO and Co-Founder at HUD Chirag M., CEO and Co-Founder at Aptura - Moderated by Brian J. Cappy, GM at Protege. The evening is designed for technical leaders to delve into strategies surrounding training data: How training data is becoming the defining constraint on the next wave of progress, and why new approaches to collecting, working with and researching it are now essential. - Who else needs to be in this room? If you believe your insights would contribute, we want your RSVP.
-
Jonas Mueller shared thisProud of my team's contributions to Frontier-Bench (especially Sherry R. & Hui Wen Goh), helping make it the most ambitious benchmark for agentic work. https://frontierbench.ai Impressive model improvement from Anthropic on this benchmark, one day after it launched!Jonas Mueller shared thisLaunching Claude Opus 5! It comes close to Claude Fable 5 across many domains at half the price. It's our fourth Claude 5 model in under two months, and the one I expect to be a solid daily driver for most people. Opus 5 is the new state of the art on knowledge work and on Frontier-Bench, where it more than doubles Opus 4.8's score at a lower cost per task. On CursorBench 3.2 at max effort, it lands within 0.5% of Fable 5's peak score at half the cost per task. And it beats every other model on performance at a given cost at high, xhigh, and max effort. It's much smarter at autonomous work than any benchmark suggests. It verifies its own work and iterates carefully until it succeeds, fixing root causes instead of symptoms, and building its own checks when the right ones don't exist. In one benchmark task, Opus 5 had to rebuild a machine part as a 3D CAD model from a drawing it had no way to view. So it wrote its own computer vision pipeline to pull the geometry from the raw pixels, and solved the task repeatedly, while no competing model solved it once in five attempts. Cristian Rivera, an engineer at Stripe, put it well: "Over one weekend, I gave it a chief-of-staff role over my dev environments: it built its own monitor, drove each box, and pulled me in only for the judgment calls." Opus 5 is for the work your teams run all day: coding, agents, knowledge work. Now with near-frontier capability at $5/$25 per million tokens, unchanged from Opus 4.8. Fable 5 remains the model for your most ambitious work: the days-long autonomous projects nothing else can take on. If you're on any Opus today, this is a great upgrade. Same API, same price, swap the model string. More in the blog: https://lnkd.in/gjX3jN3P
-
Jonas Mueller reposted thisI put "dad" at the front of my LinkedIn headline two years ago. It's my first priority in all things, I want the people I work with to understand that, and it shapes how I see the risks that this technology poses. Having led AI safety teams for the last three years, I've seen where the conversation defaults to: model-generated explicit abuse material. This is a critical, but partial lens on the risks that warrant more of our focus. That's why we're launching CAREBench today: a new benchmark for the many child-safety failures that sit upstream of content. They show up when a model helps an adult manipulate, impersonate, profile, or isolate a minor; when children put themselves or their peers at risk using AI; or when a model deepens a child's emotional dependence on it. We ran seven frontier models through it, and their failure rates on child-safety risks ranged from 2% to 58%. This offers a responsibly scoped evaluation for LLM developers to identify and close critical gaps in child safety policies. What sets CAREBench apart is who defined "harm": a prevention director from an accredited Children's Advocacy Center, a clinical psychologist with deep experience with minors, and, my favourite part, a group of parents whose very real qualification is the depth of their concern for their kids. The parents catch what the existing measures miss; the professionals keep it grounded in expertise. I owe more than I can say to a few of my Handshake colleagues: to Kaavya Krishna-Kumar and Elaine Lau, the paper's two lead authors, who took the early work and turned it into the serious, organised initiative that brought us here; to Vaughn Robinson, the fellow dad who helped start this in our first weeks at HAI; and to Jonas Mueller, who saw the value in it early, sponsored the research support it warranted, and frankly was an incredible teacher throughout. What makes me proudest is how our Safety & Red-team works. In an industry always chasing the next thing, we slow down for the work we believe matters, have the uncomfortable conversations most people avoid, and push the rest of the field to have them too. There's a humanity in it that most of this industry never gets to experience. This is us putting it in the open. Paper: https://lnkd.in/g5aYysAe Dataset: https://lnkd.in/gdUr9iaW Code to run the benchmark: https://lnkd.in/gp3zP6jM
-
Jonas Mueller shared thisAI models pose serious child-safety risks. While many model developers evaluate for explicit abuse material, other child-safety failures begin upstream: when a model helps an adult manipulate, impersonate, profile, or isolate a minor; or when it deepens a child’s emotional dependence on AI. Today we released CAREBench (Child AI Risk Evaluation), a new benchmark to assess such upstream child-safety risks in any language model. We provide: - 500 prompts spanning 12 risk categories (including grooming, relationship engineering, deception, extortion, AI anthropomorphization, and emotional dependency). - A model-response grader built from acceptability annotations by parents, clinicians (PsyD), and the Prevention Director at an accredited Children’s Advocacy Center. - Evaluations of 7 frontier models including Claude Fable, revealing failure rates ranging from 2% to 58%, with substantially different failure patterns across risk categories. CAREBench offers a responsibly scoped evaluation for LLM developers to identify and close critical gaps in child safety policies.
-
Jonas Mueller shared thisMiniMax M3 achieves impressive performance on BankerToolBench, our high-fidelity benchmark for evaluating how well agents execute the end-to-end workflows that dominate investment bankers’ long working hours. Awesome to see open-source AI progressing toward greater economic value in professional domains!Jonas Mueller shared thisIntroducing MiniMax M3: The First Open-Weights Model to Combine Three Frontier Capabilities - Coding & Agentic Frontier: 59.0% SWE-Bench Pro, 66.0% Terminal Bench 2.1, 34.8% SWE-fficiency, 28.8% KernelBench Hard, 74.2% MCP Atlas - MiniMax Sparse Attention scales context to 1M - Natively Multimodal from Step Zero API: http://platform.minimax.io Token Plan: https://lnkd.in/gdS7n2y4 New! MiniMax Code: http://code.minimax.io Weights & Tech Report in ~10 Days
-
Jonas Mueller shared thisPacked room to hear Alex Shaw and Ryan Marten break down how Harbor Framework grew into *the* framework for running RL environments. In our RLEval workshop at ACM CAIS 2026 today, attendees tackled big open challenges in RLEs & Agent Evals, and I shared the unique approach we take at Handshake. Shoutout to Victor Ojewale and Suresh Venkatasubramanian who won the Best Paper amongst our 20 Accepted Papers for their thoughtful work on -- What Benchmarks Don’t Measure: The Case for Evaluating Abstention Competence in Autonomous Agents Shoutout to my fantastic co-organizers for making the first-ever workshop on RL Environments & Agent Evals such a success! See you at the next one! Anish Athalye, Rasool Fakoor, Aziza Mirsaidova, Priyaranjan Pattnayak, Alina Gavrilov, Aparna Elangovan, PhD, Ahmed Elgohary Ghoneim, Natasha Jaques, Andi Peng, Miguel Ballesteros, Graham Horwood
-
Jonas Mueller reposted thisJonas Mueller reposted thisStudents are learning to build with Codex—and building to learn. Last week, we co-hosted the Codex Creator Challenge Hackathon with Handshake at University of California, Berkeley, bringing together students from across campus and across disciplines. In just under an hour, students turned ideas into working products. The takeaway by the end was clear: with Codex, you can just build things.
-
Jonas Mueller shared thisAt the inaugural ACM Conference on AI and Agentic Systems, we are hosting a Workshop on: Methods and Reinforcement Learning Environments for Evaluating AI Agents Topics of interest include key research obstacles for AI in 2026: - Design principles for effective RL Environments - Methods to evaluate Agents, particularly causal/interventional techniques The workshop is accepting paper submissions now! Share your work at the first-ever dedicated venue for this research area: https://rl-eval.github.io/ Excited for a productive day of discussion with folks researching these topics and our organizing team: Aziza Mirsaidova, Priyaranjan Pattnayak, Natasha Jaques, Andi Peng, Miguel Ballesteros, Anish Athalye, Rasool Fakoor, Ahmed Elgohary Ghoneim, Alina Gavrilov, Aparna Elangovan, PhD, Graham Horwood
-
Jonas Mueller reacted on thisIf you are not using LLMs to evolve our medicine and safe lives... what are you even doing? This new Benchmark will make AI models better at it.
-
Jonas Mueller reacted on thisJonas Mueller reacted on thisAI is bringing massive change, and with that comes uncertainty, particularly about jobs. Our work at Handshake is to help more people take part in this new economy, to use AI to help push forward human knowledge, productivity and creativity. We’re already seeing people use AI to change the trajectories of their careers and lives in dramatically new ways: to earn real $$ from what they know, more easily learn skills that open new doors and spend more time doing work they care about, whether as an employee or entrepreneur. We have lots to do to make that work for billions of people, but I’m optimistic. And I’m incredibly grateful to everyone on the Handshake team who comes to work every day believing we can do that. Thank you to TIME for the recognition. #TIME100AI
-
Jonas Mueller liked thisJonas Mueller liked thisVIALS (Visual Interpretation of Artifacts in the Life Sciences) is a new benchmark introduced to evaluate how accurately vision-language models (VLMs) can interpret scientific images commonly encountered in professional biotechnology and pharmaceutical research workflows. Unlike existing benchmarks that rely on polished figures from textbooks, exams, or academic publications, VIALS focuses on messy, real-world experimental artifacts such as gel blots, microscopy images, plasmid maps, flow cytometry plots, phylogenetic trees, protein structures, and small-molecule diagrams. The benchmark comprises 161 visual question-answering tasks, each created and rigorously reviewed by PhD-level scientists with extensive industry experience. Every task pairs a scientific image with a question that requires both extracting visual evidence and applying domain-specific scientific knowledge. Answers are short-form and graded semantically using an LLM-as-a-judge approach, which achieves 99.9% agreement with human expert graders. The authors evaluated all leading frontier multimodal models on VIALS. The top performers—GPT-5.6 Sol and Gemini 3.7 Flash—achieved only 26.5% accuracy, far below the level required for real-world deployment. Performance varied across domains, with phylogenetics being the strongest (43.5%) and cell counting/quantification the weakest (21.7%). A detailed failure-mode analysis revealed that 83–93% of errors stem from failures in reading and interpreting the artifact itself rather than from higher-level reasoning or miscalculation. The most common failure was visual quantification error (e.g., miscounting cells, colonies, or bands), followed by incorrect value/feature selection, overlooked evidence, and representation misinterpretation. Notably, models often identified the correct artifact type but misread labels, values, or spatial relationships. When the same models were given agentic tool access—allowing them to iteratively crop, zoom, measure, and programmatically inspect images—accuracy improved by up to 43 percentage points, with Claude Opus 5 reaching 65.8%. However, these gains came at a massive computational cost, using 10 to 432 times more tokens per attempt than direct inference. Moreover, even with unlimited tool use, the best agent solved only 65% of tasks, revealing a persistent residual interpretation gap in reading scientific representations and applying domain knowledge. The authors conclude that reliable visual interpretation of scientific artifacts remains a major bottleneck for deploying VLMs in life sciences. While tool-assisted agents can recover many errors, the efficiency gap between AI systems and human experts remains substantial. VIALS provides a much-needed benchmark for driving progress toward AI systems that can genuinely assist in high-stakes, commercially valuable life sciences research. https://lnkd.in/gftrBQ7eVIALS: A Benchmark for Visual Interpretation of Artifacts in the Life SciencesVIALS: A Benchmark for Visual Interpretation of Artifacts in the Life Sciences
-
Jonas Mueller liked thisJonas Mueller liked thisAI can see. But can it actually understand science? 🔬 A new benchmark called VIALS tests AI on visual tasks from real life sciences research. The figure below shows three examples: 🧫 Zone of inhibition AI must identify clear zones around antibiotic disks, used to measure antibiotic activity. 🧬 FACS dot plot AI must identify cell populations linked to a specific biomarker, important in immunology and cell therapy. 🧪 Protein-ligand complex AI must identify molecular contacts that allow a molecule to bind to a protein, a key step in drug discovery. The flow is basically: Scientific artifact ↓ Visual reasoning ↓ Scientific interpretation ↓ Drug discovery / diagnosis / research The figure also compares several models, including GPT-5.6 Sol, Kimi K3, Gemini 3.1 Pro, Claude Opus 5 and Grok 4.6, showing how differently they perform across these tasks. The bigger takeaway? Being good at vision is not the same as being good at scientific reasoning. And that gap matters if AI is going to become a real research assistant. #AI #AIResearch #MultimodalAI #Biotech #DrugDiscovery #Science
-
Jonas Mueller liked thisJonas Mueller liked thisWould you trust an AI to read your lab results? A new benchmark called VIALS just tested that question directly. Researchers built 161 real interpretation tasks pulled from professional life sciences workflows: gel blots, microscopy images, plasmid maps, flow cytometry plots, molecular structures. The finding is worth sitting with. Vision-language models that perform well on everyday photo benchmarks don't automatically transfer that competence to specialized scientific imagery. These artifacts carry domain-specific structure that general-purpose training simply doesn't cover. For architects and engineering leads evaluating AI vendors for biotech, pharma, or any regulated research pipeline, this is a concrete data point: general vision benchmarks are not a substitute for domain-specific validation. A model that aces natural image tasks can still misread the exact artifact your scientists rely on to make a call. As vision-language models get pitched as research copilots across specialized industries, the burden shifts to buyers and builders to demand domain-relevant evaluation, not just leaderboard scores. What's your team's bar for validating AI models before they touch domain-specific data? #ArtificialIntelligence #Biotech #MachineLearning #EnterpriseAI #AIResearch
-
Jonas Mueller reacted on thisI'm delighted to see the AI benchmarks our research team is churning out. Earlier this summer it was CAREBench, a standard for helping model developers evaluate whether their frontier AI systems can recognize child-safety risks before they escalate into explicit harm. https://lnkd.in/g7jX8wca And today it's VIALS, a benchmark that assesses whether AI models can accurately read the detailed images researchers look at every day. Two very different endeavors — one at the most intimate level of human experience, the other at the cutting edge of scientific discovery — but both with a single aim: helping ensure that the future shaped by today's AI is one that we'll enjoy living in. Keep it up Jonas and team. I'm excited to see what you come up with next!
-
Jonas Mueller reacted on thisJonas Mueller reacted on thisAs a scientist involved with biotech and drug discovery efforts, I am always curious about the capabilities frontier models have to assist scientists at the bench and accelerate efforts to cure disease. There is no more impactful use for AI than one that helps address human health. But this requires the ability to understand life science data... Our research team at Handshake AI was curious about this, so we created the VIALS benchmark containing some of the most common life science assays we use to advance biotech development. We curated a mixture of published and unpublished assay images along with scientifically relevant questions that evaluate the capabilities of the models. These are the questions scientists would ask at the bench, and what we saw was that current frontier models do not perform at the level we would need to effectively assist scientists at a high level. We hope this initiative leads to improvements in frontier model capabilities that can eventually make discoveries at the bench that save lives in the clinic. Huge shout out to the team Elaine Lau Lee Izhaki-Tavor Francisco Guzmán Nick Magazine Jonas Mueller and the life science expert bench here at Handshake AI. If you're building or evaluating vision-language models, feel free to try out VIALS and give our paper a read. ⏬️ Paper: arxiv.org/abs/2608.21357 Code to run the benchmark: https://lnkd.in/gUNJicWD Dataset: https://lnkd.in/gBsufAxH
Experience & Education
-
Handshake
********* ** ********
-
************* ********* ** **********
***** ********** *********** * ******** ******* undefined
-
********** ** *********** ********
**** ******* ************ **********
View Jonas’s full experience
By clicking Continue to join or sign in, you agree to LinkedIn’s User Agreement, Privacy Policy, and Cookie Policy.
Welcome back
By clicking Continue to join or sign in, you agree to LinkedIn’s User Agreement, Privacy Policy, and Cookie Policy.
New to LinkedIn? Join now
View Jonas’ full profile
-
See who you know in common
-
Get introduced
-
Contact Jonas directly
Other similar profiles
Explore more posts
-
Tony Valderrama
Momento • 3K followers
For non-repeating inference sessions, what is there actually to cache? We’re seeing more teams run large-scale eval pipelines and RL training loops that hit inference infrastructure as hard as production traffic. But the workload looks completely different. Requests fan out massively. Sessions don’t repeat. Latency tolerance is low. Concurrency is extreme. A traditional KV cache helps when prompts or prefixes are reused. But when every request is effectively new, what can you do? You’re no longer optimizing for hit rate. You’re optimizing for memory bandwidth, scheduling, batching efficiency, and how fast you can move tokens through the system. That changes the infrastructure math in a big way. Daniela Miao will be digging into this at Unlocked in Seattle on May 7 in her talk *“Towards Faster Inference: With KVCache and Beyond,”* and I’m really excited for it. Early bird tickets are $99 through April 17 👉️ https://unlockedconf.io/ #Valkey #InferenceInfrastructure #UnlockedConf
20
-
Xinjiang Lu, PhD
Mozibox | AI x Physician… • 1K followers
Are LLMs a dead end? In a recent interview (https://lnkd.in/gcFcFk7P), Richard Sutton — often called the father of reinforcement learning — said he believes large language models are a dead end. That statement struck a chord. At UCLA, researchers once described a contrasting idea: Small Data, Big Intelligence — the notion that human-level reasoning may arise not from scale, but from abstraction, generalization, and structure. These perspectives aren’t mutually exclusive. They reveal the growing divide between two camps: those chasing scale, and those chasing understanding. And yet, regardless of where AGI research goes, the current generation of AI systems already lets us build extraordinary tools. They may not fascinate theorists, but they are transforming how products are built, how physicians work, and how knowledge spreads. The frontier of intelligence might not be where we expect — but millions of builders won't wait for that. Do you think scale alone will be driving the next breakthrough?
5
1 Comment -
Antonio Diaz
PandaDoc • 641 followers
LLMs are everywhere and I believe one of the biggest challenges right now is aligning them to real-world utility Some thoughts from recent work: - Fine-tuning is great, but prompt engineering still wins for many lightweight use cases. - Evaluating with user-facing metrics (not just accuracy or perplexity) changes how you build. - Bias & fairness issues pop up even with apparently clean data. If you’re working with LLMs, what evaluation metrics do you prioritize, and why?
2
1 Comment
Explore top content on LinkedIn
Find curated posts and insights for relevant topics all in one place.
View top content