Nobody upstairs cares which model you picked or how elegant the graph looks. Four things stand between a working prototype and budget approval: predictable cost, evidence, recovery, no lock-in. That's the whole list. https://lnkd.in/ge9fuWgE #AIAgents #Infrastructure Sam Yang
Diagrid
Software Development
Making AI agents and Workflows reliable and secure in production.
About us
Diagrid's mission is Make AI agents reliable and secure. Out product Diagrid Catalyst is an agentic AI orchestration platform unifying agent durability and governance, designed for Kubernetes. Diagrid Catalyst works with any AI agent framework and MCP server to add resilient, self-recovering execution, zero-trust security, and cryptographic identity to every agent and tool. Built on open-source standards, Catalyst automatically resumes agents and workflows from failures, reduces downtime, protects critical resources with policy-based access control, lowers LLM and infrastructure costs through reliable execution, and accelerates the path from prototype to production.
- Website
-
https://diagrid.io
External link for Diagrid
- Industry
- Software Development
- Company size
- 51-200 employees
- Headquarters
- Seattle
- Type
- Privately Held
- Founded
- 2021
- Specialties
- AI Agent Infrastructure, Durable Execution, Durable Workflows, AI Agent Reliability, Production AI Agents, Multi-Agent Systems, Workflow Orchestration, Agent Observability, Zero-Trust Security for AI, MCP Servers, Distributed Systems, Dapr, Enterprise AI Platform, Platform Engineering, Cloud-Native Workflows, Microservices, Pub/Sub Messaging, Service Discovery, Event-Driven Architecture, and Kubernetes
Products
Diagrid Catalyst
DevOps Software
Diagrid Catalyst is an agentic AI orchestration platform unifying agent durability and governance, designed for Kubernetes. Diagrid Catalyst works with any AI agent framework and MCP server to add resilient, self-recovering execution, zero-trust security, and cryptographic identity to every agent and tool. Built on open-source standards, Catalyst automatically resumes agents and workflows from failures, reduces downtime, protects critical resources with policy-based access control, lowers LLM and infrastructure costs through reliable execution, and accelerates the path from prototype to production.
Employees at Diagrid
Locations
-
Primary
Get directions
Seattle, US
Updates
-
The LangGraph checkpointer is good. There's one failure mode it doesn't cover: the pod dies after the API call lands but before the checkpoint write. Now you have two Jira tickets and two charges. 🎩 Marc Duiker on keeping LangGraph and fixing that seam. https://lnkd.in/g-iqSUhr #LangGraph #AIAgents #DevOps
-
LangGraph, CrewAI, Strands, ADK, Microsoft Agent Framework. Take off the branding and they're all the same five-line loop. Which means durability doesn't belong inside the framework. Mark Fussell on why coupling the two is how the industry gets locked in for a decade. https://lnkd.in/gQsCJCC7 #AIAgents #SoftwareArchitecture
-
Diagrid reposted this
Struggling with durable workflows & agents? Try out our new plugin: I'll be using Claude but if you're using Codex or Copilot, find the docs for install and login here: https://lnkd.in/eCR6v62N -- claude plugin marketplace add diagridio/catalyst-ai && claude plugin install catalyst-ai@diagrid -- Then connect by running /mcp -> Authenticate. Fire away the prompt: -- I'm new to Diagrid Catalyst. Show me a durable workflow that survives a crash, using the Catalyst MCP tools for everything in Catalyst, and your shell for git and running the app. 1. Call `catalyst_whoami` and tell me which org I'm signed in to. If that fails, help me sign in. 2. Clone https://lnkd.in/e-h8FVgZ, take the durable workflow sample in my language (Python if unsure), and make sure the app `durable-workflow` exists in project `default`. 3. Get its connection with `catalyst_get_connection`, write the values to a `.env` in the sample folder (gitignored), and start the sample in the background with its environment loaded from it. 4. Run one workflow to completion. Then start one through the sample's crash endpoint, show me with `catalyst_get_workflow_run` that it's RUNNING and waiting, not failed, and wait until I type continue. 5. Restart the app. When the run completes, show me from its history that finished steps ran once and only the step the crash cut off ran again. 6. Explain in at most three sentences why it survived, then offer one next step and wait for my pick. End with this run's console link: https://lnkd.in/eMQ3RVyh<app-id>/<instance-id>?project=<project>. -- It's that easy to get started! You can also write your own prompt and get creative. Catalyst is free to get started with, give it a spin and lmk what you think 👀
-
Every vendor is "agentic" now. Same movie as "cloud" a decade ago, except these things touch customer data and money. Three questions that expose who's faking it: kill the demo mid-task, ask for the record of what the agent did, ask what a failure costs. https://lnkd.in/gstC3-yK #AIAgents #AgenticAI Sam Yang
-
Diagrid reposted this
Learning about Durable agents with Dapr Agents. Great insights by Yaron Schneider and Mark Fussell at Agentic AI Night. Lu.ma/aaif-sea-02
-
Diagrid reposted this
If you're running the local Dapr Dev Dashboard, please update it (just restart, it will auto update) type in the Konami code (↑ ↑ ↓ ↓ ← → ← → B A) and have some fun! (let me know your highscore) Download instructions: https://lnkd.in/e7eETU8K
-
👀 ✍ Let agent frameworks define the harness. Let the runtime make it survive production.
A "universal agent harness" on top of LangGraph, OpenAI Agents, ADK, Strands, etc. is just middleware cosplay. You take an existing framework, wrap it, then recreate: sessions, tools, approvals, events, subagents, streaming. Congrats, now developers have two frameworks to worry about, leaky abstractions, more bloat and extra failure points. This is exactly why we took the opposite approach with Dapr: We don’t wrap the framework and re-invent the AI primitives. Instead, we make the infrastructure underneath it durable and secure. Keep LangGraph. Keep OpenAI Agents. Keep Strands. Keep whatever wins next month. Because here’s what always happens with yet-another-harnesses: Either the abstraction leaks, or the underlying frameworks get dumbed down until they fit it - a race to the bottom where you lose either way. The framework should own how the agent is built: its loop, tools, handoffs, subagents, and semantics. The underlying runtime should own durability, retries, state, messaging, identity, mTLS, observability, and recovery. That separation matters, because the moment you put a "universal harness" in the middle, you’re betting that this abstraction can keep up with every agent framework without leaking or flattening their differences. That’s a bad bet. Let agent frameworks define the harness. Let the runtime make it survive production.
-
Second place to start in Dapr Ops Dashboard: the Cluster View. Every Dapr cluster in one console, plus a live map of how your apps actually call each other. Teams routinely find a dependency nobody documented, usually while looking for something else. https://hubs.ly/Q04ywG420 #Dapr #Kubernetes #Observability
-