# braintrust.dev > AI-optimized mirror of braintrust.dev containing 50 pages totalling 58,884 words of clean markdown content, structured data, and semantic HTML. Original source: https://braintrust.dev. Last updated: 2026-07-20T14:25:12.821Z. Each page is available as HTML (with JSON-LD structured data) and Markdown (text-only, ideal for LLMs and RAG). ## Homepage - [Braintrust - The AI observability platform for building quality AI products](/content/site-root.html): Ship quality AI at scale. Braintrust is the AI observability platform for tracing production, running evals, and catching regressions before they reach users. (373 words) ## Articles & Blog Posts - [docs/integrations/sdk-integrations/opentelemetry/index.html](/content/docs/integrations/sdk-integrations/opentelemetry/index.html) (986 words) - [Coda's Help Desk with and without RAG - Braintrust](/content/docs/cookbook/recipes/codahelpdesk/index.html) (2,101 words) - [Generating release notes and hill-climbing to improve them - Braintrust](/content/docs/cookbook/recipes/releasenotes/index.html) (1,903 words) - [Autoevals - Braintrust](/content/docs/sdks/typescript/related/autoevals/overview/index.html): Autoevals is a tool to quickly and easily evaluate AI model outputs (1,571 words) - [Best practices - Braintrust blog - Braintrust](/content/blog/best-practices/index.html): Browse all Best practices articles from the Braintrust team. (444 words) - [Create datasets - Braintrust](/content/docs/annotate/datasets/create/index.html): Build datasets from CSV uploads, the SDK, production logs, user feedback, traces, or Loop. Includes multimodal records like images and attachments. (1,446 words) - [Unreleased AI: A full stack Next.js app for generating changelogs - Braintrust](/content/docs/cookbook/recipes/unreleasedai/index.html) (1,898 words) - [API Reference - Braintrust](/content/docs/api-reference/index.html): Complete API reference for the Braintrust API (1,164 words) - [Product - Braintrust blog - Braintrust](/content/blog/product/index.html): Browse all Product articles from the Braintrust team. (425 words) - [Privacy Policy - Braintrust](/content/legal/privacy-policy/index.html) (1,701 words) - [How Replit identifies patterns across millions of AI sessions - Customers - Braintrust](/content/customers/replit/index.html): Learn how Replit moved from manual, multi-tool debugging to a unified observability layer that surfaces patterns at scale across its AI agent. (1,004 words) - [Authentication - Braintrust](/content/docs/admin/authentication/index.html): Understand how users and services authenticate to Braintrust, and how Braintrust authenticates to model providers on your behalf. (642 words) - [Pricing - Braintrust](/content/pricing/index.html): Start building AI products for free with Braintrust. Transparent pricing for AI evaluation, monitoring, and observability. No hidden fees, pay as you scale. (231 words) - [Security - Braintrust](/content/docs/security/index.html) (1,048 words) - [Use the Braintrust gateway - Braintrust](/content/docs/deploy/gateway/index.html): Route requests to any AI provider through a unified LLM API (491 words) - [Data Processing Addendum - Braintrust](/content/legal/dpa/index.html) (949 words) - [CI/CD integration definition - Braintrust](/content/encyclopedia/ci-cd-integration/index.html): Running eval suites automatically on pull requests or before deployments. This prevents regressions from being merged or shipped. (156 words) - [How Browserbase helps customers build reliable browser agents - Customers - Braintrust](/content/customers/browserbase/index.html): Learn how Browserbase uses evals and observability to tackle the unique challenges of building reliable browser agents at scale. (902 words) - [blog/img/gpt56-decision-map/cost_quality_interactive-html.html](/content/blog/img/gpt56-decision-map/cost_quality_interactive-html.html) (96 words) - [Observe your application - Braintrust](/content/docs/observe/index.html): Turn every production request into a searchable trace. Filter, analyze, and drill into real-time data to identify issues and gather signal for improving your AI system. (525 words) - [Datasets - Braintrust](/content/docs/annotate/datasets/index.html): Versioned collections of test cases that power repeatable evaluations and capture real production behavior as your application evolves. (134 words) - [Autoevals Python API - Braintrust](/content/docs/sdks/python/related/autoevals/0-3-0.html): Python API reference for Autoevals v0.3.0 (321 words) - [How we chose the model behind Topics with Baseten - Blog - Braintrust](/content/blog/model-behind-topics-baseten/index.html): How Braintrust and Baseten evaluated small models for Topics, using evals to improve Issues recall while keeping trace intelligence affordable at scale. (1,351 words, Jul 15, 2026) - [Evaluating speech-to-text models - Blog - Braintrust](/content/blog/voice-evals-stt/index.html): I ran a controlled eval across six speech-to-text providers, 240 audio clips, and eight content domains to find where voice agents break and which STT model comes out ahead. (1,154 words, Jul 9, 2026) - [7 best unified LLM API providers in 2026 - Articles - Braintrust](/content/articles/best-unified-llm-api-providers-2026/index.html): Compare top unified LLM API providers: Braintrust Gateway, OpenRouter, Vercel AI Gateway, LiteLLM, Portkey, Together AI, and Groq. (2,883 words, Jun 25, 2026) - [5 best AI agent observability tools for agent reliability in 2026 - Articles - Braintrust](/content/articles/best-ai-agent-observability-tools-2026/index.html): Compare the top AI agent observability platforms: Braintrust, Agenta, Fiddler, Helicone, and Galileo for production agent monitoring and evaluation. (1,615 words, Jun 21, 2026) - [Best LLM tracing tools for multi-agent systems (2026 review) - Articles - Braintrust](/content/articles/best-llm-tracing-tools-2026/index.html): Compare the leading LLM tracing tools: Braintrust, Arize Phoenix, Langfuse, Galileo AI, Maxim AI, Fiddler, and Helicone. (2,378 words, Jun 21, 2026) - [Best RAG observability tools (2026): monitor retrieval and generation in production - Articles - Braintrust](/content/articles/best-rag-observability-tools-2026/index.html): Compare Braintrust, Arize Phoenix, Langfuse, Comet (Opik), and Galileo for production RAG observability across retrieval tracing, live quality scoring, drift detection, and self-host options. (1,180 words, Jun 21, 2026) - [How to test agent cost-efficiency with Braintrust - Blog - Braintrust](/content/blog/test-agent-cost-efficiency/index.html): How the cheapest model is not always the cheapest system, and how evals plus control logic reduce cost per resolved request. (1,632 words, Jun 17, 2026) - [Best AI agent analytics tools (2026): see trends across every agent answer - Articles - Braintrust](/content/articles/best-ai-agent-analytics-tools-2026/index.html): Compare Braintrust, Galileo, HoneyHive, Datadog LLM Observability, and Langfuse across answer classification, multi-facet filtering, source-trace review, evaluation handoff, and pricing at production scale. (1,299 words, Jun 12, 2026) - [How to track LLM costs (2026): A playbook for per-user, per-feature, and per-agent-run attribution - Articles - Braintrust](/content/articles/how-to-track-llm-costs-2026/index.html): A 2026 playbook for tracking LLM costs at the request level, with patterns for per-user, per-feature, and per-agent-run attribution, normalized provider pricing, and release gates that protect quality. (3,256 words, Jun 2, 2026) - [Best AI governance platforms for LLM applications (2026): Eval, audit, and enforce - Articles - Braintrust](/content/articles/best-ai-governance-platforms-llm-applications-2026/index.html): Compare the best AI governance platforms for LLM applications in 2026. See how Braintrust, Galileo, Credo AI, Fiddler AI, and Patronus AI cover eval-time scoring, production audit, RBAC, and runtime enforcement. (3,365 words, May 29, 2026) - [Braintrust CLI and MCP - Blog - Braintrust](/content/blog/cli-and-mcp/index.html): Learn when to use the Braintrust CLI and MCP depending on where you are in the AI development workflow. (746 words, Apr 3, 2026) - [How to build your first offline eval - Blog - Braintrust](/content/blog/offline-eval-guide/index.html): A 10-step guide to going from a vibe to a working eval system, using a real Mermaid diagram generation project as an example. (3,056 words, Mar 10, 2026) - [Best Promptfoo alternatives in 2026: Open-source tools and SaaS - Articles - Braintrust](/content/articles/best-promptfoo-alternatives-2026/index.html): Compare the best Promptfoo alternatives for LLM evaluation in 2026. See how Braintrust, DeepEval, and RAGAS compare for production AI evaluation, team collaboration, and CI/CD integration. (2,954 words, Mar 3, 2026) - [Trace keynote recap: See it, improve it, optimize it - Blog - Braintrust](/content/blog/trace-keynote/index.html): Everything we announced at the Trace keynote, including Topics, the Braintrust CLI, and the Braintrust Gateway. (1,056 words, Feb 25, 2026) - [Braintrust's series B: building the infrastructure for production AI - Blog - Braintrust](/content/blog/announcing-series-b/index.html): Braintrust has raised $80M to become the observability layer for shipping quality AI. (685 words, Feb 17, 2026) - [7 best AI observability platforms for LLMs in 2025 - Articles - Braintrust](/content/articles/best-ai-observability-platforms-2025/index.html): Compare the top AI observability platforms: Braintrust, Langfuse, Galileo AI, Helicone, Maxim AI, Fiddler AI, and Evidently AI. (1,305 words, Dec 19, 2025) - [Braintrust Java SDK: AI observability and evals for the JVM - Blog - Braintrust](/content/blog/java-sdk/index.html): AI observability and evaluation tools for Java applications, built on OpenTelemetry. (713 words, Oct 23, 2025) - [Top 10 LLM observability tools: Complete guide for 2025 - Articles - Braintrust](/content/articles/top-10-llm-observability-tools-2025/index.html): Compare the leading LLM observability platforms for production AI applications. (3,351 words, Oct 2, 2025) - [AI that knows your data - Blog - Braintrust](/content/blog/mcp/index.html): Introducing Braintrust's MCP server. (465 words, Sep 9, 2025) - [GPT-5 vs. Claude Opus 4.1 - Blog - Braintrust](/content/blog/gpt-5-vs-claude-opus/index.html): Which one you should ship with, and how to know for sure. (755 words, Aug 8, 2025) - [Brainstore is now on by default - Blog - Braintrust](/content/blog/brainstore-default/index.html): Brainstore is now the default in both our UI and API. Learn what's changing and coming next. (669 words, Mar 31, 2025) - [Bedrock, Vertex AI, and universal structured outputs - Blog - Braintrust](/content/blog/model-updates/index.html): Full support for Bedrock, Vertex AI, and structured outputs in the AI proxy and playground. (366 words, Feb 14, 2025) - [Custom scoring functions in the Braintrust Playground - Blog - Braintrust](/content/blog/custom-scorers/index.html): Create custom scorers and access them via the Braintrust UI and API., In this video, I demonstrate the new online scoring feature in Braintrust for evaluating use cases. By configuring custom scores in the playground, we can assess the performance of outputs like jokes based on themes. No specific action is requested from viewers. (489 words, Sep 16, 2024) - [Braintrust achieves SOC 2 Type II compliance - Blog - Braintrust](/content/blog/soc2/index.html): We are excited to announce that Braintrust has achieved SOC 2 Type II compliance. (106 words, Jul 15, 2024) - [Weekly update 10/09/23 - Blog - Braintrust](/content/blog/update-1/index.html): Performance improvements, fine tuning tutorial, Alpaca Evals, autocomplete in the playground. (272 words, Oct 9, 2023) - [It's time to build reliable AI - Blog - Braintrust](/content/blog/reliable-ai/index.html): Introducing Braintrust: the enterprise-grade stack for building AI products. From evaluations, to prompt playground, to data management, we take uncertainty and tedium out of incorporating AI into your business. (1,207 words, Sep 12, 2023) ## About Pages - [Contact us - Braintrust](/content/contact/index.html): Get in touch with the Braintrust team. Talk to our sales team about enterprise solutions or join our Discord community for support and discussions about AI development. (65 words) ## Resources - [Full Page Index](/index.html): Browse all cached pages with rich metadata - [About This Cache](/content/about.html): Methodology, technical details, and usage guidelines - [XML Sitemap](/sitemap.xml): Machine-readable sitemap for crawler discovery - [Robots.txt](/robots.txt): Crawler directives