Overview
klaw is designed to handle failures gracefully. This page covers the resilience mechanisms built into the provider layer, context management, structured error handling, and observability.Provider Retry
When an LLM API call fails with a retryable error, klaw automatically retries with exponential backoff and jitter.Retry Flow
Backoff Calculation
Example progression (3 retries):
- Attempt 1: ~1s wait (1s + jitter)
- Attempt 2: ~2s wait (2s + jitter)
- Attempt 3: ~4s wait (4s + jitter)
Retryable Errors
All other errors (e.g., HTTP 400, 401, 403) fail immediately without retry.
Fallback Chain
When retries are exhausted on the primary provider, klaw tries fallback providers in order:Configuration
Context Window Compaction
When conversation history approaches the context window limit, klaw automatically compacts it via LLM-based summarization.Compaction Lifecycle
Configuration
Trigger point: Compaction occurs when estimated tokens exceed
(200,000 - 8,192) * 0.75 ≈ 143,856 tokens.
Token estimation uses a ~4 characters per token heuristic. The summarized section is prefixed with [Previous conversation summary].
Structured Errors
All agent errors use theAgentError type with machine-readable codes:
Error Codes
Error Format
Errors serialize as[CODE] message: cause:
Unwrap() for Go error chain inspection.
Observability
Structured Logging
klaw uses Go’sslog package for structured JSON logging:
Metrics
Global and per-session metrics are tracked using atomic counters: Global Metrics:
Per-Session Metrics:
Metrics are recorded via:
RecordRequest(sessionID, inputTokens, outputTokens)— after each LLM callRecordToolCall(sessionID, toolName)— after each tool executionRecordError(sessionID, errorCode)— on any error
metrics.Summary() to get a snapshot of all global counters.
Next Steps
Agent Loop
How these mechanisms fit into the execution cycle
Cost & Safety Guide
Practical guide to configuring safety features

