Langfuse’s cover photo
Langfuse

Langfuse

Software Development

Open Source LLM Engineering Platform, now part of ClickHouse

About us

Langfuse is an open source AI engineering platform. It helps teams collaboratively develop, monitor, evaluate, and debug AI applications. Langfuse can be self-hosted in minutes and is battle-tested and used in production by thousands of users from YC startups to large companies like Khan Academy or Twilio. Langfuse builds on a proven track record of reliability and performance. Developers can trace any Large Language model or framework using our SDKs for Python and JS/TS, our open API or our native integrations (OpenAI, Langchain, Llama-Index, Vercel AI SDK). Beyond tracing, developers use Langfuse Prompt Management, its open APIs, and testing and evaluation pipelines to improve the quality of their applications. Product managers can analyze, evaluate, and debug AI products by accessing detailed metrics on costs, latencies, and user feedback in the Langfuse Dashboard. They can bring humans in the loop by setting up annotation workflows for human labelers to score their application. Langfuse can also be used to monitor security risks through security framework and evaluation pipelines. Langfuse enables non-technical team members to iterate on prompts and model configurations directly within the Langfuse UI or use the Langfuse Playground for fast prompt testing. Langfuse is open source and we are proud to have a fantastic community on GitHub and Discord that provides help and feedback. Do get in touch with us! Langfuse is now part of ClickHouse.

Website
https://langfuse.com
Industry
Software Development
Company size
11-50 employees
Headquarters
San Francisco
Type
Privately Held
Founded
2022
Specialties
Langfuse, Large Language Models, Observability, Prompt Management, Evaluations, Testing, Open Source, LLM, AI, Analytics, Open Source, and Artificial Intelligence

Products

Employees at Langfuse

View 30 employees at Langfuse

or

By clicking Continue to join or sign in, you agree to LinkedIn’s User Agreement, Privacy Policy, and Cookie Policy.

See all employees

Locations

Updates

  • View organization page for Langfuse

    15,738 followers

    We're running through New York City! 🗽 Two weeks ago, 20 engineers ran in Paris. Next week, we'll go again. this time around AI Engineer NYC. Join us for a relaxed 5km sightseeing run through central park followed by free breakfast and coffee on us + meet some Langfuse faces. Sign up in the comments ⬇️

    • No alternative text description for this image
    • No alternative text description for this image
  • Banks need evidence before an LLM reaches production. Collecting it by hand takes weeks, and by then the prompt or model has already changed. Together with one of the world's five largest banks, we scripted that evidence: golden financial datasets, domain-specific evaluators, and a PASS/FAIL gate for every model, prompt, and agent. Every verdict links back to traces in Langfuse that a reviewer can inspect. Built by Doneyli De Jesus. Full write-up and open-source repo in the comments.

  • Learn how Langfuse helps AI to fix itself in Marc's talk 👇

    In June I gave a talk at AI Engineer World's Fair SF on how the best teams we work with at Langfuse are now self-improving their agents. The loop is the same as before: trace production, build datasets, write evals, experiment, ship. What's new is who drives it. Agents now propose fixes and hill-climb against your evals. Over the last month we've also seen teams let agents suggest new dataset items and evaluators based on errors they find in production. My take: don't hand off everything. Fully automated loops burn a lot of tokens and produce a lot of slop. Stay involved at the top, deciding what goes into the dataset, what counts as failure, and which changes ship, and let agents handle the tedious parts. You spend less time, and quality goes up, because an agent can read far more traces than you ever will. This also changes the data layer. Langfuse used to be mostly write-heavy. Now agents read and query a lot more of it, and traces from a year ago are useful context. So you want to own that data and keep all of it, unsampled. Link to full recording in comments!

    • No alternative text description for this image
  • Langfuse is coming to AI Engineer NYC, Oct 12–14. Two sessions to meet us: → 𝗟𝗮𝗻𝗴𝗳𝘂𝘀𝗲 𝗪𝗼𝗿𝗸𝘀𝗵𝗼𝗽: 𝘁𝗵𝗲 𝗔𝗜 𝗲𝗻𝗴𝗶𝗻𝗲𝗲𝗿𝗶𝗻𝗴 𝗹𝗼𝗼𝗽, 𝗲𝗻𝗱 𝘁𝗼 𝗲𝗻𝗱 Mon, Oct 12 · 2:20pm · Sutton (Leadership Track) → 𝗕𝘂𝗶𝗹𝗱𝗶𝗻𝗴 𝗗𝗲𝗽𝗹𝗼𝘆𝗺𝗲𝗻𝘁 𝗚𝗮𝘁𝗲𝘀 𝗳𝗼𝗿 𝗔𝗜 𝗔𝗴𝗲𝗻𝘁𝘀 𝗶𝗻 𝗙𝗶𝗻𝗮𝗻𝗰𝗶𝗮𝗹 𝗦𝗲𝗿𝘃𝗶𝗰𝗲𝘀 𝘄𝗶𝘁𝗵 Doneyli De Jesus Wed, Oct 14 · 2:55pm · McCarthy (Curated) The rest of the time, find us at the booth on the Expo floor for demos and questions about tracing, evals, or running agents in regulated environments. Schedule in the comments.

  • OpenAI's decision API is now available for evals in Langfuse Use it to score, classify and catch signals on your production traces. OpenAI's decision API is an alternative to TypeSafe AI's Jev entering the market. 𝗪𝗵𝘆 𝘂𝘀𝗲 𝗮 𝗱𝗲𝗰𝗶𝘀𝗶𝗼𝗻-𝗼𝗿𝗶𝗲𝗻𝘁𝗲𝗱 𝗺𝗼𝗱𝗲𝗹? ✔️ 𝗖𝗵𝗲𝗮𝗽𝗲𝗿. Both Jev and gpt-6-luna are priced by input tokens only. TypeSafe reports that Jev is 40 to 400x cheaper than frontier models on classification tasks. OpenAI prices gpt-6-luna at $0.10 per million input tokens, with no output-token charge. ✔️ 𝗙𝗮𝘀𝘁𝗲𝗿. Decision models skip free-form generation. TypeSafe reports that Jev is 20 to 200x faster than comparable LLMs, while OpenAI reports that the Decisions API is about 10x faster than the Responses API. ✔️ 𝗖𝗮𝗹𝗶𝗯𝗿𝗮𝘁𝗲𝗱. Yes / no answers return a probability, while Choice and Score answers include a probability distribution and confidence value where available. ✔️ 𝗖𝗼𝗻𝘀𝗶𝘀𝘁𝗲𝗻𝘁. Decision models return typed values from a fixed answer space instead of sampling a free-form response. This makes them a good fit for regression tests and criteria you track over time. Check full changelog in comments 👇️

  • Langfuse reposted this

    Why did our usage jump last week? is a question we kept hearing, and until now it was hard to answer at the org level. 💸 So I built a Usage breakdown chart in Langfuse. 📊 It shows your units (traces, observations, and scores) as stacked bars across every project in your organization: → By project: spot which project is behind a spike, then jump straight into it → By type: see what your usage is made of, e.g. when an evaluator adds more scores than you expected → Download as CSV: one row per time bucket, project, and type, handy for splitting costs across teams Changelog in comment.

    • No alternative text description for this image
  • Langfuse goes Brazil 🇧🇷 This Thursday, Felipe from iFood will share his experience with Langfuse and ClickHouse at our in-person meetup in São Paulo. Para a nossa comunidade brasileira 🇧🇷: ansiosos para encontrar vocês lá. Link de inscrição nos comentários. Sign up in comments 🔽

    View organization page for ClickHouse

    157,422 followers

    Observabilidade de agentes em escala: o caso do iFood com Langfuse. Como saber o que seus agentes de IA realmente fizeram, onde acertaram e onde falharam? No dia 8 de outubro, em São Paulo, Felipe Rodrigues Garé , Sr. AI Engineer no iFood, mostra como o time usa ClickHouse e Langfuse para rodar agentes de IA sobre dezenas de terabytes de logs. Investigações que levavam uma semana de trabalho de um analista hoje terminam em cerca de duas horas. Quinta-feira, 8 de outubro, das 18h às 21h, em São Paulo. Gratuito, vagas limitadas. Solicite participação: https://lnkd.in/eewCB442 No mesmo dia, das 9h às 17h, temos o treinamento gratuito Observabilidade com ClickStack: laboratórios práticos com ClickHouse, OpenTelemetry e HyperDX, com certificação no final. Reserve seu lugar: https://lnkd.in/eDWwKngz #ClickHouse #Langfuse #Observabilidade #AgentesDeIA #SaoPaulo

    • No alternative text description for this image
  • 🇺🇸 Event alert: On October 15, Marc Klingen will speak at Deutsche Bank in North Carolina about Langfuse and how to bring agents to production. We'll cover: - Trace agents in production so you can see what's happening inside them - Learn from failures and find out why an agent went wrong - Turn real interactions into evaluation datasets that reflect how users behave - Keep improving prompts, models, tools, and agent architectures Sign up in comments.

  • Lotte Verheyden joined experts from LangChain, Mercor, CoreWeave, and Galileo to discuss what evals actually look like in production: from online vs. offline evaluation and sampling costs to why LLM judges shouldn’t use 1-10 scales. A practical conversation on building evals from real user behavior. Thanks Basil Chatha and AngelList for bringing it together!

    Evals are the most important part of building AI systems, but there’s not a lot out there about building them well. So I hosted a fireside chat on evals with Liam Bush (LangChain), Lotte Verheyden (Langfuse), Braden Holstege (Mercor), Emmanuel Turlay (CoreWeave), and Soumya Mohan (Galileo) to walk me through the tricks of the trade - how they actually work in production, where teams get them wrong, and where they're headed over the next 12 months. We get into: -online vs offline evals -why running evals can cost you more than running the actual agent -what nobody can see after the agent makes a tool call -why understanding accounting rules might be harder than AGI -and much, much, more… Some of my favorite parts: 1/ Running your evals can cost more than running your agent. "We deployed an agent where we had a 100% sample rate initially, and it 4x’d the budget, because we're spending 3x as much tokens running the live evals as we're actually running the agent." Mercor’s team caught this on one of their own deployments. If you're running live evals on every user input, you're paying for a second agent whose entire job is watching the first one - that gets expensive, really fast. Over time, they reduce the sampling rate to reduce the cost, especially as the agent improves. 2/ “Ship WITHOUT evals.” "I can't believe I have to say that, because I would say the opposite to people a few years ago. But for an agent that is not safety critical or not regulated, just launch and let it fail and learn from that, and build your evals from there." You can guess how users are going to use your product all you want, but you won’t actually know until you launch and observe them using it. Broad online evals catch the stuff you'd never have tested for, and you convert that into offline evals afterward. He did add one caveat - don't let it fail for too long or your users will churn. 3/ Never ask an LLM to rate anything 1 to 10. "It's tempting to ask the LLM judge to score on a scale of zero to ten. This is really bad. You should probably not do that at all." Just like its hard for humans to know the difference between a 5 and a 6, its hard for models too. Classifiers are cheaper and more accurate. Example: if you're grading politeness on call transcripts, your options should be insufficiently polite, sufficiently polite, and extremely polite: "I don't want my agent to be extremely polite. That's actually a bad case. I don't want overly wordy responses, or the user will be frustrated if something is not working and the LLM is like, 'you're absolutely right' - because we all hate that response." A 1-10 scale gives that response a 10. Categories force you to decide what you actually want, which is the harder and more useful exercise. Link to full episode in the comments. — Follow me Basil Chatha for everything AI Agents. And thanks to AngelList for making this possible!

Similar pages

Browse jobs