For AI agents: a documentation index is available at the root level at /llms.txt. Append /llms.txt to any URL for a page-level index, or .md for the markdown version of any page.
PII redaction: added organization-level redaction settings with separate controls for Gateway requests and stored telemetry, configurable categories, sensitivity controls, and previews.
Model leaderboard: added a Top models by spend chart to help users understand Gateway usage on the landing page.
Observability: combined Tracing and Observability into a single experience while preserving existing links and workflows.
Simulation: added a streamlined simulation flow where users can select a prompt and version directly, generate scenarios, run simulations, and inspect results.
Gateway: BYOK requests can now fall back to managed credits when provider credentials encounter rate limits.
Model leaderboard: legends now reflect the entire selected time range, model colors remain consistent, and individual models can be isolated for comparison.
PII detection: improved redaction reliability for short telemetry fields, Unicode content, street addresses, private IPv4 addresses, and loopback IPv6 addresses.
Tables: the first column now remains consistently frozen, preventing alignment from shifting across views.
Behavior classification: improved processing reliability so behavior results remain current during traffic spikes.
Exports: partially completed exports now provide the rows successfully written, identify incomplete time ranges, and preserve failure reasons for later investigation. JSONL exports are also written as JSONL rather than JSON, which could cause downstream systems to count records twice.
PII settings: fixed saved settings losing an explicitly selected organization or the default category configuration.
Slack integration: fixed large channel lists causing setup requests to time out.
Saved filters: invalid filter requests now return actionable validation errors instead of server errors.
Security: strengthened organization access controls and prevented duplicate cash-credit events from changing balances more than once.
Models: improved model-detail layouts, search relevance, and provider matching. Model selectors also refresh after custom model or provider updates without clearing active searches and filters. Deployment-specific models now appear correctly when configuring BYOK provider credentials.
Gateway: provider timeouts now return a clearer timeout status, and connection-level provider failures return more accurate error responses.
Gateway: improved Anthropic passthrough fidelity for errors, streaming responses, optional fields, and retry behavior.
Responses API: OpenRouter models can now be called through the Responses API.
Logs: object parameters such as metadata now appear as individual fields in the Config tab instead of a raw JSON object.
Demo workspace: updated sample prompts and added introductory guidance throughout the demo experience.
Settings: improved page rendering and migrated organization data loading to a more reliable query system.
Tables: numeric values are now consistently right-aligned, with improved header spacing across product tables.
Experiments: fixed columns becoming hidden after returning to the experiment list from a detail page.
Scores: fixed errors when opening evaluation-result details backed by PostgreSQL.
Logs: fixed the first and last time labels being clipped on events charts, and fixed an issue that prevented users from clearing the selected grouping.
Models: fixed incorrect model filters being applied when reloading model-detail pages.
Monitors and reports: corrected breakdown values and removed duplicate data requests.
Behaviors: added LLM judging to improve behavior classification and analysis.
Experiments: experiment details now present grader names, outcomes, scores, latency, cost, tokens, and errors more clearly while hiding unnecessary metadata.
Experiment performance: added early row previews, lighter list responses, and compression for large experiment APIs so results appear sooner and transfer less data.
Model discovery: improved recommended-model ordering, BYOK availability filtering, and model-detail layouts to make suitable models easier to find and compare.
Model catalog: compressed model catalog responses to reduce loading time.
Workflows: optimized workflow execution paths for substantially faster processing.
Monitors: simplified monitor rendering for clearer, more consistent notifications.
Log exports: large exports now use stable ordering, recover from storage interruptions, and avoid silently dropping rows. Failed or incomplete exports also report accurate progress and no longer provide untrustworthy downloads.
Reliability: brief cache interruptions no longer restart otherwise healthy web servers or turn into avoidable request failures.
Models: plain and provider-prefixed model names now resolve to the same catalog entry and behave consistently.
Spans: model icons remain readable when the Model column is narrow.
Lists: invalid sorting options now fall back safely instead of causing server errors.
Dashboard metrics: fixed organization-filtered quantile requests that could return server errors.
Evaluator results: LLM graders now provide a written rationale alongside each score, making it easier to understand why an evaluation passed or failed.
Experiments: refined evaluation analytics and comparison views while reducing unnecessary trace loading, and added backend status visibility and automatic refreshes for active experiments.
Evaluator editor: added expanded editors for LLM prompts and Python code, with immediate synchronization and smoother keyboard interactions.
Error Tracking: improved how errors are classified into incidents for more accurate investigation.
Model discovery: models now default to newest-first ordering, with clearer release dates and lifecycle information.
Model pages: improved layouts across compact, expanded, sparse, and data-rich views.
Experiment performance: compressed large experiment-log responses to reduce loading time and transferred data.
Hosted MCP: improved sign-in continuity for hosted MCP connections with a dedicated OAuth session flow.
Logs API: removed internal platform fields from log, trace, and span API responses.
Reports: reports can now include Pulse summaries for richer incident and behavior reporting.
Logs: request activity can now be grouped by Threads or Traces directly from Logs, with filtering, sorting, refreshing, and detail panels available in each grouped view.
Experiments: experiment lists now include evaluator scores and model details, making results easier to compare without opening each experiment.
Behaviors: Behaviors is now available to regular platform users, with a complete draft, training, and review workflow.
Home: added a recently visited section for quickly returning to prompts and datasets.
Red Team: improved campaign setup, engine performance, CLI presentation, documentation, and error handling.
Experiments: conditional evaluator results can now return N/A. Aggregate scores exclude unscored rows and show the number of scored results, while histograms display N/A separately.
Evaluators: refined the grader editor, expanded Code evaluation access, and improved how LLM and Code configurations are displayed and saved.
Behaviors: improved classification accuracy, reduced false positives, added controls for flagging incorrect matches, and made behavior charts clearer across time ranges.
Error tracking: improved error classification and added automatic clustering for errors occurring around the same time.
Model catalog: improved model information and catalog performance, including richer descriptions, release dates, and coding benchmark scores.
Request logs: customer names and email addresses are now preserved in request-log lists, and nearby-request navigation more accurately follows the selected event’s timestamp.
Dashboard sessions: bookmarked dashboards now refresh expired sessions automatically instead of unexpectedly returning users to sign-in.
Metrics: empty breakdown charts now remain in a stable no-data state without flashing.
Prompts: new drafts are created only after an edit, keeping prompt version history cleaner.
Dataset imports: imported message histories now appear as conversations, and expected outputs are stored and displayed as formatted JSON.