Skip to main content

Overview

The agent loop is the heart of klaw. It orchestrates the conversation between user input, LLM decisions, and tool execution, with built-in safety guardrails for iteration limits, context management, cost tracking, and human approval.

The Loop

Step-by-Step Flow

1

Receive Input

User sends a message through a channel (CLI, Slack, API).
2

Planning Phase (Optional)

On the first message, if planning is enabled, a planning prompt is injected asking the LLM to outline steps before acting.
The default planning prompt asks the LLM to:
  1. Analyze what the user is asking for
  2. List steps (max 5)
  3. Identify potential issues
3

Context Window Check

Before each LLM call, the context manager estimates token usage. If the conversation exceeds the compaction threshold, middle messages are summarized via an LLM call.
  • Max context: 200,000 tokens (default)
  • Compaction threshold: 75% of available context
  • Reserve: 8,192 tokens for the response
  • Keeps first user message and last 6 messages verbatim
4

Send to LLM

Agent sends conversation history to the LLM provider with resilient delivery (retry + fallback).
5

Process Stream

Agent collects streaming events: text chunks, tool calls, stop signals with usage data.
6

Human Approval (Optional)

If a tool is in the require_approval list, the user is prompted before execution.
The user sees: ⚠ Tool 'bash' requires approval. Execute? [y/N]:
7

Execute Tools (Parallel)

Approved tools execute concurrently. Each tool runs in its own goroutine with a 2-minute timeout. Results are collected in original call order after all goroutines complete.
If the LLM returns 3 file reads simultaneously, they complete in roughly the time of one read instead of three sequential reads.
8

Reflection (Optional)

After every N tool calls (default: 3), a reflection prompt is injected asking the LLM to assess progress.
9

Cost & Iteration Check

The loop checks two safety limits before continuing:
10

Loop or Complete

If tools were called, loop back to the context check step. If no tools were requested, the response is complete.

Iteration Limit

Each agent has a maximum number of loop iterations (default: 50). This prevents runaway loops where the LLM repeatedly calls tools without converging on a solution.
When the limit is reached, the agent stops with AgentError{Code: ErrMaxIterations}.

Context Window Management

Large conversations can exceed the LLM’s context window. The context manager handles this automatically: Compaction strategy:
  1. Keep the first user message verbatim (preserves original intent)
  2. Keep the last 6 messages verbatim (preserves recent context)
  3. Summarize everything in between via an LLM call
  4. Replace middle messages with [Previous conversation summary]
Token estimation uses a ~4 characters per token heuristic.

Cost Tracking

Every LLM call records input and output token counts. The cost tracker computes session cost using per-model pricing tables: Set a budget with max_session_cost in config. When the budget is reached, the agent stops with ErrBudgetExceed. A warning is logged at 80% of the budget.

Structured Errors

The agent loop uses typed errors with machine-readable codes:
Errors format as [CODE] message: cause and support errors.Unwrap().

Stream Events

The provider returns a stream of events: The stop event now includes a Usage field with token counts, used by the cost tracker.

Tool Execution

Timeout

Each tool has a 2-minute timeout by default:

Result Formatting

Tool results are formatted for display:

Error Handling

Tool errors are captured and fed back to the LLM as structured errors:

Memory Integration

Before each LLM call, workspace context is loaded:

Sub-Agent Delegation

The delegate tool allows the main agent to spawn ephemeral sub-agents that run inline during a conversation. This is useful for parallelizable sub-tasks or specialized work.
Key properties:
  • Sub-agents use RunOnce — a non-streaming agent loop with parallel tool execution
  • Each sub-agent gets its own tool registry (parent delegate excluded, child delegate added at depth+1)
  • Maximum nesting depth: 3 levels
  • 5-minute timeout per delegation
  • Output truncated at 30,000 characters

Performance Considerations

Next Steps

Resilience

Provider retry, fallback, and error handling

Cost & Safety

Budgets, approval, and monitoring

Tools Reference

Available built-in tools

Custom Tools

Build your own tools