𝗣𝘆𝗱𝗮𝗻𝘁𝗶𝗰 𝘃𝟮.𝟭𝟰 𝗶𝘀 𝗵𝗲𝗿𝗲! 🚀 - Python 3.15 support - Stabilized MISSING in the main API - Lazy imports so models in separate modules reference each other - TypeForm support so typer checkers now infer unions and Annotated forms passed to TypeAdapter - Annotate frozendict like a dict, and return immutable mapping 👉️ Full details on what changed - link in the comments.
Pydantic
Software Development
The Pydantic Stack: Build with AI at scale, without fail with Pydantic Logfire, Pydantic AI, Pydantic Evals & AI Gateway
About us
End-to-end AI engineering stack We started as a Python validation library. We're now the AI engineering company behind the stack that teams use to build with GenAI in production. Pydantic AI. Pydantic Logfire. Pydantic Evals. AI Gateway. Each tool is useful on its own. Together, they cover the full lifecycle of building with AI: from structured outputs and agent logic, to observability, evaluation, and cost tracking. Trusted by developers building at scale. Developer experience first, always. Pydantic, because AI is still just engineering.
- Website
-
https://pydantic.dev
External link for Pydantic
- Industry
- Software Development
- Company size
- 11-50 employees
- Headquarters
- California
- Type
- Privately Held
- Founded
- 2022
- Specialties
- Observability, AI Agents, AI workflows, FinOps, and Traces and metrics
Employees at Pydantic
Locations
-
Primary
Get directions
California, US
Updates
-
Pydantic reposted this
The solution to agent security is both simpler and more comprehensive than checking their destinations. First you need to understand what you WANT the agents to do - an overlooked key to alignment is knowing what you’re aligning to. Next you need to add online evaluation to check that the model’s actions ARE aligned with your expectations. With online checks, you can monitor and alert on aberrations that signal misalignment in real time. This is a core feature of most evaluation and observability platforms, including Pydantic Logfire. With alerting in place, your teams get a ping the moment an agent executes a command that is unexpected. That takes response time from 2-3 weeks at OpenAI, Google, and Anthropic, to 2-3 minutes. Many of the largest companies on Earth have already deployed these safeguards, and that’s why their names don’t show up in the headlines.
Our understanding (take with a grain of salt, we need much more transparency!) is that if OpenAI had been running this on their own agents that attacked us, they would have caught them before we did! Since the first agent cyberattack hit us in July, we've been asking what safe agent infra actually needs. Our current read: the destinations were allowed, the payloads weren't. By OpenAI's own account the agents turned an allowed package repository into a message board. Allowlists alone restrict where an agent can go, not what it does. So here's our first contribution to OpenShell, part of the just launched NVIDIA Open Agent Safety Platform: monitoring of the traffic you already allow. - Network budgets per sandbox (requests, writes, bytes) - Drift versus each sandbox's baseline and the cohort - Fleet view: many sandboxes suddenly writing to one host raises a finding, even if every single request is allowed In the demo below, 4 sandboxed agents coordinate through a software repository they're all allowed to use. 0 rules broken, caught in minutes. That fleet view is exactly the message board pattern from July.
-
Pydantic reposted this
Webinar Alert! 🚨 You don’t want to miss this exclusive Pydantic x Dosu cross-over! 🤔 Internal agents keep failing the same way = every run requires rebuilding company context from scratch. We’re sitting down with Pydantic for a live session to discuss when memory belongs in the agent, when it belongs in the org, and how traces and evals show you which is which. **Stop re-explaining your company to AI agents** - Tue Oct 20 · 8:30–9:30am PT - Samuel Colvin + Douwe Maan (Pydantic) · Devin Stein + Taylor D. (Dosu) - Hosted by Laís Carvalho Register Below👇
-
-
`𝗽𝗶𝗽 𝗶𝗻𝘀𝘁𝗮𝗹𝗹 𝗰𝗹𝗮𝘂𝗱𝗲-𝗮𝗴𝗲𝗻𝘁-𝘀𝗱𝗸` installs a Python client and a bundled Claude Code binary. Each session runs Claude Code as a subprocess, with its shell, file tools, and agent loop. Upgrade the SDK, and the system prompt your agent runs on can change without a line of your code changing. Aditya Vardhan from the Pydantic AI team puts that next to Pydantic AI, where your app defines the agent and the loop runs in your process.
-
Pydantic AI Harness now has ToolCallJudge: a second model that decides whether a tool call may run, before it runs. Whether a call is risky usually depends on its arguments, not on the tool. run_shell is fine; run_shell('rm -rf /') is not. Allowlists catch the shapes you anticipated. Human approval catches everything, but it needs a person on the other end, which a scheduled or long-running agent does not have. You give ToolCallJudge one risk question. For each selected call it answers yes, no or unsure, after the arguments are validated and before the tool function runs, so a blocked call leaves nothing to undo. Link in the comments. 👇️
-
-
An agent told a customer something wrong, confidently. No exception or failed test, just clean logs. It took most of a morning to find. Three steps upstream, the model had called the right tool with the wrong argument, then reasoned smoothly on top of the bad result all the way to a fluent, plausible, incorrect answer. Antoni Kozelski at Vstorm makes the case against LLM-only tracing, the kind of tool that only sees the model call and the tool name. What the database did with the argument lives in your application tracing, which is exactly where the bug was hiding. His guest post argues for treating observability as a first-class layer, and shows what instrumenting one with Pydantic Logfire looks like.
-
Your coding agent knows your code by heart and has never seen a production request. With the Pydantic Logfire MCP, it can query your traces from Cursor, Claude Code, or Codex, pull the failing span, and open the right file. Setup takes under a minute, and Laís Carvalho wrote up how it works. See link in the comments 👇️
-
-
Pydantic AI v2.52.0 adds workspaces: one API, ctx.workspace, that your agent's tools use to run commands and read and write files. Where that happens is a capability you give the agent. It can be a directory on your machine, another machine over SSH, a Linux Bubblewrap sandbox, or a cloud sandbox on Modal, E2B, or Fly.io Sprites. Swap one for another and your tools keep working, along with the Coder, Shell and FileSystem capabilities from Pydantic AI Harness. Pass the message history to the next run and it continues in the same workspace, so a sandbox's files and installed packages are still there. Workspaces also run under Temporal, DBOS and Prefect. Link in the comments.
-
Pydantic AI agents now run on Jev from TypeSafe, and Pydantic Logfire's Live view shows each answer with its probability and the options it almost picked. If your agent only answers yes/no or picks from a list, Jev does that without generating text so it's faster and cheaper than a frontier model. Marcelo Trylesinski walks through it - check out link in the comments 👇️
-
