Investigate error incidents
The Errors page groups failed spans into incidents and error groups. Use it to measure impact, find the failures behind an outage, and inspect matching events.
If Errors does not appear in the navigation, contact your Respan administrator or Respan support.
What you can answer
- When did an error spike start, and is it still ongoing?
- Which providers, models, endpoints, or customers were affected?
- Which error groups contributed most?
- Has someone acknowledged and investigated the incident?
Read the error timeline
Errors uses two levels of grouping:
The streamgraph shows error volume across the selected range. Colored layers represent the ten largest error groups; the rest are combined into Other errors. Hover a legend item or group row to isolate its layer.
Incident rows show acknowledgement, state, severity, error count, and start time. Hover a row to highlight its window, then select it for details. Error-group rows show an aligned occurrence trend and Unresolved, Resolved, or Ignored state.
Incidents are detected from error-rate behavior. They are separate from monitors configured by your organization.
Set the scope
Choose the environment and time range
Include the incident and enough time before it to show the normal baseline.
Investigate an incident
Select an incident to open its detail panel. The header summarizes its error rate, environment, duration, affected scope, and change from the normal rate. Tags show severity, dominant error type, state, acknowledgement, and whether impact is estimated.
Use the detail sections to:
- Rank Contributing error groups by share and occurrence count. Select one to drill into it while retaining a breadcrumb to the incident.
- Review affected providers, models, endpoints, and customers.
- Compare failed and total requests, normal and peak error rates, start and end times, trigger, fault domain, and latest Data through timestamp.
- Select Acknowledge when someone takes ownership.
- Add investigation context under Internal notes.
Internal notes are shared. Do not include secrets, credentials, customer content, or raw prompts.
Ongoing incidents refresh every minute while the page is visible. Check Data through before treating current counts as final.
Investigate an error group
Open a group from the timeline or an incident’s contributor list. A contributor drill-down scopes headline counts to the incident; sampled events continue to use the page’s time range.
The Overview tab shows occurrences, diagnostics, fingerprint metadata, and a representative event payload. Use the status control to Resolve, Ignore, or Reopen the group.
Open Events to inspect sampled occurrences. Break them down by Model, Customer, or Environment, select values to narrow the list, then select an event for its log detail.
Events are sampled, not an exhaustive export. Use the occurrence total for impact and Logs for broader record-level analysis.
Triage an incident
- Choose the affected environment and a range with a short baseline.
- Check the incident severity, error rate, duration, and Data through value.
- Review contributing groups and affected dimensions.
- Acknowledge the incident and add an ownership note.
- Open the dominant group and inspect its diagnostics and events.
- Continue in Logs or Traces to confirm the first failing span and root cause.
- Resolve handled groups and create a monitor for a measurable recurring condition.
Choose the right view
Troubleshooting
The timeline has no incidents
Widen the range, clear filters, and confirm the environment. Error groups can appear without crossing the threshold for an incident; check Logs if both are empty.
Filtering hides unexpected data
Fault domain applies to incidents and groups. Error type, provider, and HTTP status affect the groups, not the incident list. Clear filters, then add them one at a time.
An expected event is missing
The Events tab is sampled. Open Logs with the same environment, time range, fingerprint, provider, and HTTP status, then search for the event or request ID.
Related guides
Compare errors with traffic, latency, cost, and cohort breakdowns.
Inspect individual spans and apply record-level filters.
Reconstruct the workflow around a failed span.
Alert on error count or error-rate thresholds.
Need help?
Join our Discord — we’ll help you investigate an error incident.