A Racket Coding Agent
The source code for this example is in the directory coding-agent-harness.
The Agentic Loop
Modern large language models are not limited to answering questions in a single turn. When given access to tools (i.e., callable functions that can read files, run commands, or search the web) an LLM can operate as an autonomous agent: it reasons about what it needs to know, calls a tool to gather information, receives the result, and continues reasoning until the task is complete. This pattern is called an agentic loop.
For a coding assistant, the loop typically looks like this:
- The user describes a change or a bug to fix.
- The LLM decides it needs to read a file and calls
read_file. - After seeing the file contents, the LLM proposes an edit via
propose_edit. - The user reviews the colored diff and approves or rejects it.
- On approval, the agent writes the file and runs
make check. - The LLM reads the check result and either continues or summarizes what changed.
The key architectural insight is that the LLM is stateless between API calls. It only knows what is in the message history. The agent accumulates tool results into that history turn by turn, giving the model the context it needs to decide what to do next.
This chapter builds a complete Racket implementation of such a coding agent. The agent classifies each user request as a coding task, a general question, or a hybrid, and routes it accordingly. It supports live web search via Brave or Exa AI, renders colored unified diffs before any file is written, and gates every accepted edit on make check. Model access is configurable: named provider profiles in a JSON config file select the endpoint, the model, the generation parameters, and the pricing, so the same loop runs against either a cloud provider such as Fireworks AI or a local MLX server. The same program also works as a Unix-style command that runs a single prompt and exits, while the interactive REPL adds readline editing, persistent history, and Tab completion.
Module Architecture
The project is organized into nine source files, each with a single clear responsibility:
1 agent.rkt REPL, CLI, intent classifier, provider dispatch, context management
2 harness-config.rkt Hierarchical JSON config: providers, generation parameters, pricing
3 fireworks-ai.rkt Fireworks and OpenAI-compatible client (SSE streaming), usage, cost
4 mlx-serve.rkt Local MLX client (mlx_lm.server /v1/chat/completions), usage
5 chat-loop.rkt Provider-agnostic agentic tool loop shared by both backends
6 tools.rkt Tool registry, five coding tools, the propose_edit approval gate
7 approval.rkt Colored diff printer, y/n/s prompt
8 search.rkt Brave Search and Exa AI search backends
9 line-input.rkt Readline-backed line editing, history, and Tab completion
The dependency graph is acyclic and close to linear. harness-config.rkt reads JSON and nothing else; approval.rkt only prints and prompts; tools.rkt requires approval.rkt; chat-loop.rkt requires tools.rkt; and the two provider clients require chat-loop.rkt, tools.rkt, and harness-config.rkt. agent.rkt requires everything and owns the session state.
The one place where a static require is not enough is line-input.rkt. Racket’s readline collection raises at instantiation time when no Editline or GNU Readline shared library is installed, so an unconditional require would take the whole harness down on machines without one. The module loads readline/readline with dynamic-require inside an exception handler instead, so a missing library degrades to a plain read-line rather than a startup failure. That is also why the Makefile builds the standalone executable with ++lib readline/readline: the flag embeds the module for raco exe while leaving it uninstantiated until the REPL actually needs it.
The Provider-Agnostic Agentic Loop
The heart of the agent is a loop that is deliberately independent of any particular model vendor. It lives in chat-loop.rkt and is parameterized by a post-fn argument: a function that takes an OpenAI-style chat-completions payload and returns a normalized response hash. Both the Fireworks client and the MLX client supply their own post-fn, and the loop never talks to a network directly.
Here is the complete file:
1 #lang racket
2
3 (require racket/string
4 "tools.rkt")
5
6 (provide chat*
7 chat-with-tools*)
8
9 ;; ---------------------------------------------------------------------------
10 ;; Helpers
11
12 (define (response-message data)
13 ;; Extract the assistant message hash from a normalized response, or #f if
14 ;; the response has no usable choices/message.
15 (define choices (hash-ref data 'choices #f))
16 (and (pair? choices)
17 (let ([choice (first choices)])
18 (and (hash? choice) (hash-ref choice 'message #f)))))
19
20 (define (msg-content msg)
21 ;; The assistant message's 'content, coerced to a string. Some OpenAI-compatible
22 ;; servers (e.g. mlx_lm.server) emit "content": null on tool-only responses;
23 ;; hash-ref with a default returns the literal 'null in that case, so coerce
24 ;; any non-string value to "" here.
25 (define c (and msg (hash-ref msg 'content "")))
26 (if (string? c) c ""))
27
28 (define (without-dangling msgs)
29 (if (and (not (null? msgs))
30 (hash-has-key? (last msgs) 'tool_calls))
31 (drop-right msgs 1)
32 msgs))
33
34 ;; ---------------------------------------------------------------------------
35 ;; chat* : post-fn (listof hash) ... -> string
36
37 ;; Build a request body, omitting generation parameters the active provider
38 ;; profile did not declare. No max_tokens/temperature default is compiled in
39 ;; here; the provider config is the only source for them.
40 (define (request-payload model-id max-tokens temperature messages)
41 (define base (hash 'model model-id 'messages messages))
42 (define with-max (if max-tokens (hash-set base 'max_tokens max-tokens) base))
43 (if temperature (hash-set with-max 'temperature temperature) with-max))
44
45 (define (chat* post-fn messages
46 #:model-id model-id
47 #:max-tokens max-tokens
48 #:temperature temperature)
49 (define payload
50 (request-payload model-id max-tokens temperature messages))
51 (define data (post-fn payload))
52 (define msg (response-message data))
53 (define content (msg-content msg))
54 (if (and (string? content) (not (string=? content "")))
55 content
56 "No response content"))
57
58 ;; ---------------------------------------------------------------------------
59 ;; chat-with-tools* : post-fn (listof hash) (listof string) ... -> (values string (listof hash))
60 ;; Multi-turn agentic loop. Returns two values: final-text and final-messages.
61
62 (define (chat-with-tools* post-fn messages tools
63 #:model-id model-id
64 #:max-tokens max-tokens
65 #:temperature temperature
66 #:max-iterations max-iterations)
67 (define tools-rendered (render-tools tools))
68 (define current-messages (box messages))
69 ;; Repetition detection: track the last few tool-call signatures so a model
70 ;; stuck issuing the identical failing call is stopped early rather than
71 ;; burning all max-iterations. A signature is (name . args-json) per call,
72 ;; sorted, so multi-call batches compare as a set.
73 (define recent-signatures '())
74 (define REPEAT-WINDOW 5) ; remember the last N batches
75 (define REPEAT-LIMIT 2) ; >= 2 identical batches in the window => stuck
76
77 (define (call-signature tool-calls)
78 (sort
79 (for/list ([tc (in-list tool-calls)])
80 (define f (hash-ref tc 'function (hash)))
81 (format "~a|~a" (hash-ref f 'name "") (hash-ref f 'arguments "")))
82 string<?))
83
84 ;; Returns #t when the same batch of calls appeared >= REPEAT-LIMIT times in
85 ;; the recent window.
86 (define (seen-too-often? sig)
87 (>= (length (filter (lambda (s) (equal? s sig)) recent-signatures))
88 (sub1 REPEAT-LIMIT)))
89
90 ;; Take the rightmost (most recent) n items of a list.
91 (define (take-right lst n)
92 (if (<= (length lst) n) lst (drop lst (- (length lst) n))))
93
94 (define (append-tool-results! results)
95 ;; Append each (call-id name result-str) tuple as a tool-role message.
96 (for ([r (in-list results)])
97 (define call-id (first r))
98 (define name (second r))
99 (define result-str (third r))
100 (set-box! current-messages
101 (append (unbox current-messages)
102 (list (hash 'role "tool"
103 'tool_call_id call-id
104 'name name
105 'content result-str))))))
106
107 (define (loop iter)
108 (cond
109 [(>= iter max-iterations)
110 ;; Max iterations -- one final no-tools call for summary
111 (define payload
112 (request-payload model-id max-tokens temperature (unbox current-messages)))
113 (with-handlers ([exn:fail? (lambda (_) (values "(max tool iterations reached)" (unbox current-messages)))])
114 (define data (post-fn payload))
115 (define msg (response-message data))
116 (define content (if msg (hash-ref msg 'content "") ""))
117 (values (if (and (string? content) (not (string=? content "")))
118 content
119 "(no summary from model)")
120 (if msg
121 (append (unbox current-messages) (list msg))
122 (unbox current-messages))))]
123 [else
124 (define payload
125 (let ([base (request-payload model-id max-tokens temperature
126 (unbox current-messages))])
127 (if (null? tools-rendered)
128 base
129 (hash-set* base 'tools tools-rendered 'tool_choice "auto"))))
130 (define data (post-fn payload))
131 (define msg (response-message data))
132 (unless msg
133 (error 'chat-with-tools* "response has no 'message'. Raw: ~a" data))
134 (define tool-calls (hash-ref msg 'tool_calls #f))
135 (define content (msg-content msg))
136 ;; Append the assistant message
137 (set-box! current-messages (append (unbox current-messages) (list msg)))
138 (cond
139 [(and tool-calls
140 (list? tool-calls)
141 (pair? tool-calls)
142 (seen-too-often? (call-signature tool-calls)))
143 ;; The model is stuck re-issuing the identical call(s) -- bail out
144 ;; with an explanation instead of looping to max-iterations.
145 (values (string-append
146 "(stopped: the model repeated the identical tool call(s) "
147 (number->string REPEAT-LIMIT)
148 " times without making progress; it may be too weak for this "
149 "task or its arguments are malformed)")
150 (unbox current-messages))]
151 [(and content tool-calls (not (string=? (string-trim content) "")))
152 (displayln "")
153 (displayln (string-trim content))
154 (set! recent-signatures
155 (take-right (cons (call-signature tool-calls) recent-signatures)
156 REPEAT-WINDOW))
157 (append-tool-results! (execute-tool-calls tool-calls))
158 (loop (add1 iter))]
159 [(not tool-calls)
160 (values (or content "(empty response from model)") (unbox current-messages))]
161 [else
162 (set! recent-signatures
163 (take-right (cons (call-signature tool-calls) recent-signatures)
164 REPEAT-WINDOW))
165 (append-tool-results! (execute-tool-calls tool-calls))
166 (loop (add1 iter))])]))
167
168 (loop 0))
chat* is the single-shot path used for general questions and for the intent classifier: it builds a request, calls post-fn once, and returns the assistant’s text. chat-with-tools* is the multi-turn loop. It renders the enabled tools into OpenAI function-calling schema, appends the assistant message and then one tool-role message per result, and recurses until the model answers without asking for a tool or the iteration cap is reached.
Two details in the loop are worth calling out. First, request-payload includes max_tokens and temperature only when they are actual values. A provider profile that omits a generation parameter sends no parameter at all, which lets the server apply its own default. Nothing is defaulted in Racket code.
Second, the loop defends itself against a model that gets stuck. call-signature reduces a batch of tool calls to a sorted list of name|arguments strings, and the loop remembers the signatures of the last REPEAT-WINDOW (five) batches. If the current batch already appears in that window, seen-too-often? fires and the loop returns an explanation instead of burning the remaining iterations on the identical call. This matters most with small local models, which tend to re-issue the same malformed call when a tool result does not change their mind.
There are two edge cases worth noting. The first is a model that emits both text and tool calls in the same turn: the text is printed to the user immediately (so it can narrate “I will read the file first”), then the tools run, then the loop continues. The second is the iteration cap: at max-iterations the loop makes one final call with no tools, asking the model to summarize what it did, rather than returning nothing.
The without-dangling helper trims a trailing assistant message whose tool_calls have no matching tool results, the state a history lands in if a run is cut short. The current loop never calls it, but it is the repair to reach for if you add a path that can abandon a turn mid-flight.
The Fireworks AI Client
The API and Pricing
Fireworks AI is a hosted inference platform that serves many open-weight models through an OpenAI-compatible API. The agent does not compile in an endpoint, a model, or a price. All of those come from the active provider profile in the harness config, which is covered later in this chapter. The example profile used in this chapter declares DeepSeek Flash (accounts/fireworks/models/deepseek-v4p1-flash) at $0.14 per million uncached input tokens, $0.028 per million cached input tokens (an 80 percent cache discount), and $0.28 per million output tokens.
The estimated session cost accumulated over a conversation is:

where p is the total prompt tokens, k is the cached portion of those prompt tokens (billed at the discount), and c is the total completion tokens. The agent tracks all of these and displays the running total on demand. The rates themselves are read from the profile’s pricing block, so a different profile can declare different numbers; a profile that declares no pricing at all reports the cost as unknown rather than pretending it is free.
Streaming with Server-Sent Events
Unlike the plain one-shot chat, the Fireworks client streams its responses using server-sent events (SSE). The API sends a sequence of data: lines, each carrying a small delta of the response, terminated by a data: [DONE] line. Streaming removes any wall-clock cap on generation time, at the cost of reassembling the response on the client side. The only timeouts are CURL-MAX-TIME (seconds to wait for response headers and the TCP connection) and STREAM-IDLE-TIMEOUT (seconds of silence from the server before giving up). As long as tokens keep flowing, a request may run for minutes.
Here is the complete file:
1 #lang racket
2
3 (require net/http-easy
4 json
5 racket/string
6 racket/port
7 "tools.rkt"
8 "chat-loop.rkt"
9 "harness-config.rkt")
10
11 (provide DEBUG-LOG
12 CURL-MAX-TIME
13 active-pricing
14 prompt-cost
15 completion-cost
16 cached-cost
17 accumulate-usage
18 reset-session-stats
19 session-cost
20 print-session-stats
21 chat
22 chat-with-tools
23 make-sse-line-reader
24 parse-sse-response)
25
26 ;; ---------------------------------------------------------------------------
27 ;; Constants
28 ;;
29 ;; Endpoint, model, api_key_env, generation parameters, and pricing all come
30 ;; from the active provider profile in the harness config; no provider-specific
31 ;; value is compiled in here.
32
33 (define DEBUG-LOG (make-parameter #f))
34 ;; Requests use SSE streaming ("stream": true), so there is NO total
35 ;; wall-clock cap on generation: a long response that keeps producing
36 ;; tokens simply keeps streaming. The only remaining timeouts are:
37 ;; CURL-MAX-TIME -- seconds to wait for response headers (TTFT)
38 ;; and for the TCP connection itself.
39 ;; STREAM-IDLE-TIMEOUT -- seconds of *silence* from the server before we
40 ;; give up. Tokens arriving periodically never
41 ;; trip this; only a genuinely stalled connection
42 ;; does. Turn up / down as you like.
43 ;; http-easy's default request timeout is only 30s, so these MUST be passed
44 ;; via #:timeouts below or the constants do nothing.
45 (define CURL-MAX-TIME 600)
46 (define STREAM-IDLE-TIMEOUT 300)
47
48 ;; Pricing is read from the active provider profile's "pricing" block (USD per
49 ;; 1M tokens). A profile that declares no pricing yields #f rates, and callers
50 ;; report the cost as unknown instead of inventing a number.
51
52 (define (active-pricing)
53 (provider-pricing (config-active-provider)))
54
55 ;; ---------------------------------------------------------------------------
56 ;; Session stats (thread-safe)
57
58 (define stats-sema (make-semaphore 1))
59 (define session-prompt-tokens (box 0))
60 (define session-completion-tokens (box 0))
61 (define session-total-tokens (box 0))
62 (define session-cached-tokens (box 0))
63
64 (define (reset-session-stats)
65 (call-with-semaphore stats-sema
66 (lambda ()
67 (set-box! session-prompt-tokens 0)
68 (set-box! session-completion-tokens 0)
69 (set-box! session-total-tokens 0)
70 (set-box! session-cached-tokens 0))))
71
72 ;; Costs use the active provider's configured rates. Each returns #f when the
73 ;; profile declares no such rate, so callers can report "unknown" instead of a
74 ;; misleading $0.00.
75
76 (define (rate-cost tokens rate)
77 (and rate (* tokens rate (/ 1 1000000))))
78
79 (define (prompt-cost tokens)
80 (rate-cost tokens (pricing-ref (active-pricing) 'input)))
81
82 (define (cached-cost tokens)
83 (rate-cost tokens (pricing-ref (active-pricing) 'cached_input)))
84
85 (define (completion-cost tokens)
86 (rate-cost tokens (pricing-ref (active-pricing) 'output)))
87
88 ;; Cached input tokens are reported by the server in
89 ;; usage.prompt_tokens_details.cached_tokens and are part of prompt_tokens;
90 ;; bill them at the discounted rate and subtract them from the uncached pool.
91 (define (session-cost)
92 ;; -> number, or #f when the active profile declares no pricing at all.
93 (define rates (active-pricing))
94 (define input (pricing-ref rates 'input))
95 (define cached (pricing-ref rates 'cached_input))
96 (define output (pricing-ref rates 'output))
97 (and (or input cached output)
98 (call-with-semaphore stats-sema
99 (lambda ()
100 (define pt (unbox session-prompt-tokens))
101 (define ca (unbox session-cached-tokens))
102 (+ (or (rate-cost (max 0 (- pt ca)) input) 0)
103 (or (rate-cost ca cached) 0)
104 (or (rate-cost (unbox session-completion-tokens) output) 0))))))
105
106 (define (print-session-stats)
107 (define-values (pt ct tt ca)
108 (call-with-semaphore stats-sema
109 (lambda ()
110 (values (unbox session-prompt-tokens)
111 (unbox session-completion-tokens)
112 (unbox session-total-tokens)
113 (unbox session-cached-tokens)))))
114 (define cost (session-cost))
115 (define rates (active-pricing))
116 (displayln "")
117 (displayln "Session token usage:")
118 (displayln (format " Prompt tokens: ~a" pt))
119 (displayln (format " Completion tokens: ~a" ct))
120 (displayln (format " Total tokens: ~a" tt))
121 (when (> ca 0)
122 (define pct (* 100.0 (/ ca (max 1 pt))))
123 (displayln (format " Cached tokens: ~a (~a% of prompt)" ca (~r pct #:precision 1))))
124 (if cost
125 (displayln (format " Estimated cost: $~a ($~a/M input, $~a/M cached input, $~a/M output)"
126 (~r cost #:precision 6)
127 (~r (or (pricing-ref rates 'input) 0) #:precision 4)
128 (~r (or (pricing-ref rates 'cached_input) 0) #:precision 4)
129 (~r (or (pricing-ref rates 'output) 0) #:precision 4)))
130 (displayln " Estimated cost: n/a (no \"pricing\" block for this provider)")))
131
132 (define (accumulate-usage data)
133 (define usage (hash-ref data 'usage (hash)))
134 (when (and (hash? usage) (not (hash-empty? usage)))
135 (call-with-semaphore stats-sema
136 (lambda ()
137 (set-box! session-prompt-tokens
138 (+ (unbox session-prompt-tokens)
139 (hash-ref usage 'prompt_tokens 0)))
140 (set-box! session-completion-tokens
141 (+ (unbox session-completion-tokens)
142 (hash-ref usage 'completion_tokens 0)))
143 (set-box! session-total-tokens
144 (+ (unbox session-total-tokens)
145 (hash-ref usage 'total_tokens 0)))
146 (define details (hash-ref usage 'prompt_tokens_details (hash)))
147 (when (hash? details)
148 (set-box! session-cached-tokens
149 (+ (unbox session-cached-tokens) (hash-ref details 'cached_tokens 0))))))))
150
151 ;; ---------------------------------------------------------------------------
152 ;; API key
153 ;;
154 ;; The env var name comes from the active provider profile's api_key_env when
155 ;; a harness config is loaded; falls back to FIREWORKS_API_KEY.
156
157 (define (get-api-key)
158 (define env-name
159 (or (let ([p (config-active-provider)])
160 (and p (provider-api-key-env p)))
161 "FIREWORKS_API_KEY"))
162 (define key (getenv env-name))
163 (unless (and key (not (string=? key "")))
164 (error 'fireworks-ai "~a environment variable not set" env-name))
165 key)
166
167 ;; ---------------------------------------------------------------------------
168 ;; SSE streaming helpers
169
170 ;; Index of the first byte in `bstr` equal to byte `b`, or #f if absent.
171 (define (bytes-index-of bstr b)
172 (let loop ([i 0]
173 [len (bytes-length bstr)])
174 (cond
175 [(= i len) #f]
176 [(= (bytes-ref bstr i) b) i]
177 [else (loop (add1 i) len)])))
178
179 ;; Returns a stateful function that reads one line at a time from the SSE
180 ;; response stream `in`. Each read waits up to STREAM-IDLE-TIMEOUT seconds
181 ;; for the next byte (an idle timeout, not a wall-clock cap), so long slow
182 ;; generations never hit a total-time limit as long as tokens keep flowing.
183 ;; Each call returns a line as bytes (newline stripped) or eof at end of
184 ;; stream. Partial lines are buffered between calls.
185 (define (make-sse-line-reader in)
186 (define buf (make-bytes 4096))
187 (define acc (box #""))
188 (define (read-more!) ;; -> #t at EOF, #f after appending more bytes
189 (unless (sync/timeout STREAM-IDLE-TIMEOUT
190 (handle-evt in (lambda (_) #t)))
191 (error 'fireworks-ai
192 "stream idle timeout: no data for ~a seconds"
193 STREAM-IDLE-TIMEOUT))
194 (define n (read-bytes-avail! buf in))
195 (cond
196 [(eof-object? n) #t]
197 [else
198 (when (> n 0)
199 (set-box! acc (bytes-append (unbox acc) (subbytes buf 0 n))))
200 #f]))
201 (lambda ()
202 (let loop ()
203 (define data (unbox acc))
204 (define nl (bytes-index-of data 10)) ; 10 == #\n
205 (cond
206 [nl
207 ;; complete line available; keep the remainder for the next call
208 (set-box! acc (subbytes data (add1 nl)))
209 (subbytes data 0 nl)]
210 [(read-more!)
211 ;; EOF: whatever is left is the final unterminated line
212 (define rest (unbox acc))
213 (set-box! acc #"")
214 (if (zero? (bytes-length rest)) eof (subbytes rest 0))]
215 [else (loop)]))))
216
217 ;; Parse one SSE "data: {...}" body into a jsexpr hash (or #f on bad JSON).
218 (define (parse-sse-chunk body)
219 (with-handlers ([exn:fail? (lambda (_) #f)])
220 (string->jsexpr body)))
221
222 ;; Reconstruct the equivalent non-streaming chat-completions response hash
223 ;; from an SSE stream:
224 ;; (hash 'id ... 'model ...
225 ;; 'choices (list (hash 'message <assistant msg>
226 ;; 'finish_reason ...))
227 ;; 'usage (hash ...))
228 ;; `message` carries accumulated 'content, (deepseek) 'reasoning_content, and
229 ;; (when present) a list of 'tool_calls hashes exactly like a non-streaming
230 ;; response: each (hash 'id ... 'type "function"
231 ;; 'function (hash 'name ... 'arguments <json string>)).
232 (define (parse-sse-response in)
233 (define content-out (open-output-string))
234 (define reasoning-out (open-output-string))
235 (define tool-calls-by-index (make-hash)) ; index -> (hash 'id 'type 'name 'arguments-box)
236 (define usage #f)
237 (define finish-reason #f)
238 (define message-id (box ""))
239 (define message-model (box ""))
240 (define next-line (make-sse-line-reader in))
241 (let loop ()
242 (define line (next-line))
243 (cond
244 [(eof-object? line) (void)]
245 [else
246 ;; line is bytes: convert, then trim whitespace and trailing \r from CRLF
247 (define trimmed (string-trim (bytes->string/utf-8 line)))
248 (cond
249 [(or (string=? trimmed "")
250 (string-prefix? trimmed ":")) ; comment / keep-alive
251 (loop)]
252 [(string-prefix? trimmed "data:")
253 (define body (string-trim (substring trimmed 5)))
254 (cond
255 [(string=? body "[DONE]") (void)]
256 [else
257 (define chunk (parse-sse-chunk body))
258 (when (hash? chunk)
259 ;; API-level error inside the stream
260 (when (hash-has-key? chunk 'error)
261 (define err (hash-ref chunk 'error))
262 (define msg
263 (cond
264 [(hash? err) (hash-ref err 'message (format "~a" err))]
265 [else (format "~a" err)]))
266 (error 'fireworks-ai "API error: ~a" msg))
267 (when (hash-has-key? chunk 'id)
268 (set-box! message-id (hash-ref chunk 'id "")))
269 (when (hash-has-key? chunk 'model)
270 (set-box! message-model (hash-ref chunk 'model "")))
271 (define chunk-usage (hash-ref chunk 'usage #f))
272 (when (and chunk-usage (hash? chunk-usage))
273 (set! usage chunk-usage))
274 (for ([c (in-list (hash-ref chunk 'choices '()))])
275 (define delta (hash-ref c 'delta (hash)))
276 (define fr (hash-ref c 'finish_reason #f))
277 (when (and fr (not (equal? fr finish-reason)))
278 (set! finish-reason fr))
279 (define c-delta (hash-ref delta 'content #f))
280 (when (string? c-delta)
281 (display c-delta content-out))
282 (define r-delta (hash-ref delta 'reasoning_content #f))
283 (when (string? r-delta)
284 (display r-delta reasoning-out))
285 (define tc (hash-ref delta 'tool_calls #f))
286 (when (and tc (list? tc))
287 (for ([t (in-list tc)])
288 (define idx (hash-ref t 'index 0))
289 (define entry (hash-ref tool-calls-by-index idx #f))
290 (unless entry
291 (set! entry (make-hash (list (cons 'id "")
292 (cons 'type "function")
293 (cons 'name "")
294 (cons 'arguments (box "")))))
295 (hash-set! tool-calls-by-index idx entry))
296 (define t-id (hash-ref t 'id #f))
297 (when (and (string? t-id) (not (string=? t-id "")))
298 (hash-set! entry 'id t-id))
299 (define t-type (hash-ref t 'type #f))
300 (when (string? t-type)
301 (hash-set! entry 'type t-type))
302 (define f (hash-ref t 'function #f))
303 (when (hash? f)
304 (define f-name (hash-ref f 'name #f))
305 (when (and (string? f-name) (not (string=? f-name "")))
306 (hash-set! entry 'name f-name))
307 (define f-args (hash-ref f 'arguments #f))
308 (when (and (string? f-args) (not (string=? f-args "")))
309 (define b (hash-ref entry 'arguments))
310 (set-box! b (string-append (unbox b) f-args))))))))
311 (loop)])]
312 [else (loop)])]))
313 (define content (get-output-string content-out))
314 (define reasoning (get-output-string reasoning-out))
315 (define idxs (sort (hash-keys tool-calls-by-index) <))
316 (define tool-calls
317 (if (null? idxs)
318 #f
319 (for/list ([idx (in-list idxs)])
320 (define e (hash-ref tool-calls-by-index idx))
321 (hash 'id (hash-ref e 'id)
322 'type (hash-ref e 'type)
323 'function (hash 'name (hash-ref e 'name)
324 'arguments (unbox (hash-ref e 'arguments)))))))
325 (define message
326 (if tool-calls
327 (hash 'role "assistant"
328 'content content
329 'reasoning_content reasoning
330 'tool_calls tool-calls)
331 (hash 'role "assistant"
332 'content content
333 'reasoning_content reasoning)))
334 (hash 'id (unbox message-id)
335 'model (unbox message-model)
336 'choices (list (hash 'message message
337 'finish_reason finish-reason))
338 'usage (or usage (hash))))
339
340 ;; ---------------------------------------------------------------------------
341 ;; Low-level POST (streaming)
342
343 (define (post-fireworks payload)
344 (define api-key (get-api-key))
345 (define p (config-active-provider))
346 (define endpoint
347 (or (and p (provider-endpoint p))
348 (error 'fireworks-ai
349 "active provider profile has no \"endpoint\"; set it in the harness config")))
350 (define headers
351 (hash 'content-type "application/json"
352 'accept "application/json"
353 'authorization (string-append "Bearer " api-key)))
354 (define stream-payload
355 (hash-set* payload
356 'stream #t
357 'stream_options (hash 'include_usage #t)))
358 (when (DEBUG-LOG)
359 (displayln (format "[DEBUG] request: ~a" (jsexpr->string (hash-remove stream-payload 'messages)))))
360 (define data
361 (with-handlers ([exn:fail? (lambda (e) (error 'fireworks-ai "HTTP error: ~a" (exn-message e)))])
362 (define resp
363 (post endpoint
364 #:headers headers
365 #:json stream-payload
366 #:stream? #t
367 #:close? #f
368 #:timeouts (make-timeout-config #:request CURL-MAX-TIME
369 #:connect CURL-MAX-TIME)))
370 (define j (parse-sse-response (response-output resp)))
371 (response-close! resp)
372 (when (DEBUG-LOG)
373 (displayln (format "[DEBUG] response: ~a" (jsexpr->string j))))
374 j))
375 (when (hash-has-key? data 'error)
376 (define err (hash-ref data 'error))
377 (define msg
378 (cond
379 [(hash? err) (hash-ref err 'message (format "~a" err))]
380 [else (format "~a" err)]))
381 (error 'fireworks-ai "API error: ~a" msg))
382 (accumulate-usage data)
383 (unless (hash-has-key? data 'choices)
384 (error 'fireworks-ai "response has no 'choices'. Raw: ~a" (jsexpr->string data)))
385 data)
386
387 ;; ---------------------------------------------------------------------------
388 ;; chat / chat-with-tools -- thin wrappers over the shared provider-agnostic
389 ;; loop in chat-loop.rkt (also used by mlx-serve.rkt).
390 ;;
391 ;; Generation defaults come from the active provider profile's "generation"
392 ;; section when a harness config is loaded; explicit keyword args win.
393
394 ;; Model and generation parameters resolve from the active provider profile.
395 ;; A missing model is an error; missing generation parameters are left out of
396 ;; the request so the server's own default applies.
397
398 (define (active-model-id)
399 (define p (config-active-provider))
400 (or (and p (provider-model p))
401 (error 'fireworks-ai
402 "active provider profile has no \"model\"; set it in the harness config")))
403
404 (define (gen-param key)
405 (generation-ref (provider-generation (config-active-provider)) key #f))
406
407 (define (chat messages
408 #:model-id [model-id (active-model-id)]
409 #:max-tokens [max-tokens (gen-param 'max_tokens)]
410 #:temperature [temperature (gen-param 'temperature)])
411 (chat* post-fireworks messages
412 #:model-id model-id
413 #:max-tokens max-tokens
414 #:temperature temperature))
415
416 (define (chat-with-tools messages tools
417 #:model-id [model-id (active-model-id)]
418 #:max-tokens [max-tokens (gen-param 'max_tokens)]
419 #:temperature [temperature (gen-param 'temperature)]
420 #:max-iterations [max-iterations 20])
421 (chat-with-tools* post-fireworks messages tools
422 #:model-id model-id
423 #:max-tokens max-tokens
424 #:temperature temperature
425 #:max-iterations max-iterations))
Reassembling the SSE Stream
The SSE stream is a sequence of lines like:
1 data: {"id":"...","choices":[{"delta":{"content":"The"}}]}
2 data: {"id":"...","choices":[{"delta":{"content":" answer"}}]}
3 data: {"id":"...","choices":[{"delta":{"content":" is"}}]}
4 data: {"id":"...","choices":[{"delta":{},"finish_reason":"stop"}]}
5 data: [DONE]
make-sse-line-reader returns a stateful function that yields one line at a time, buffering partial lines between calls and applying the idle timeout. parse-sse-response then walks those lines and reassembles the deltas:
contentdeltas are appended to a string output port.reasoning_contentdeltas (for reasoning models) go to a separate port.tool_callsdeltas are the tricky part, because a single tool call’s name and arguments arrive split across many chunks. The code accumulates them in a hash keyed by the call’sindex, appending argument fragments to a boxed string. At the end it sorts the indices and rebuilds the tool-call list.- The final
usagechunk is captured for token accounting.
The output is a single normalized response hash with the same shape the non-streaming MLX backend produces, so chat-loop.rkt never knows which backend it is talking to.
Token Accounting and Cost
Because Fireworks is a paid service, the module tracks usage. The session counters live in boxes guarded by a semaphore, the same defensive pattern used for shared mutable state elsewhere in the harness. Cached input tokens are reported by Fireworks in usage.prompt_tokens_details.cached_tokens; they are part of prompt_tokens but are billed at the discounted rate. The session-cost function subtracts them from the uncached pool and bills them separately, and it returns #f when the active profile declares no rates at all. The /tokens command prints the breakdown, including the cached-token percentage.
The MLX Client
The MLX backend, mlx-serve.rkt, mirrors the Fireworks interface so that the agent loop and the REPL can swap providers with a single parameter. It targets mlx_lm.server, which exposes the OpenAI-compatible /v1/chat/completions route. That protocol already returns exactly the shape chat-loop.rkt consumes, namely choices[].message with optional tool_calls and usage.prompt_tokens / completion_tokens, so unlike Ollama’s native /api/chat there is no message re-shaping at the boundary: the module forwards the OpenAI payload as it is and returns the response unchanged.
Here is the complete file:
1 #lang racket
2
3 (require net/http-easy
4 json
5 racket/string
6 "fireworks-ai.rkt" ; for DEBUG-LOG (shared /debug toggle)
7 "chat-loop.rkt"
8 "harness-config.rkt")
9
10 (provide MLX-THINK
11 mlx-active-provider ; parameter: provider hash to read settings from
12 mlx-reset-session-stats
13 mlx-print-session-stats
14 post-mlx
15 mlx-chat
16 mlx-chat-with-tools)
17
18 ;; ---------------------------------------------------------------------------
19 ;; Constants
20 ;;
21 ;; Endpoint, model, api_key_env, and generation parameters all come from the
22 ;; active provider profile in the harness config; nothing provider-specific is
23 ;; compiled in here.
24
25 ;; The provider hash MLX requests consult for endpoint/model/generation.
26 ;; agent.rkt sets this to the active profile.
27 (define mlx-active-provider (make-parameter #f))
28
29 (define (current-provider-json)
30 (or (mlx-active-provider) (hash)))
31
32 ;; The OpenAI-compatible endpoint returns reasoning in the assistant message's
33 ;; 'reasoning field, which chat-loop.rkt ignores (it only reads 'content and
34 ;; 'tool_calls), so there is no separate thinking toggle to wire up here.
35 (define MLX-THINK (make-parameter #f))
36 ;; Non-streaming request: the whole generation must complete within this
37 ;; window. Local models on large weights can be slow, so be generous.
38 (define MLX-MAX-TIME 900)
39 (define MLX-CONNECT-TIME 10)
40
41 ;; ---------------------------------------------------------------------------
42 ;; Session stats (thread-safe). mlx_lm.server reports prompt_tokens /
43 ;; completion_tokens on every /v1/chat/completions response. Local inference is
44 ;; free, so stats are informational only -- estimated cost is always $0.
45
46 (define stats-sema (make-semaphore 1))
47 (define session-prompt-tokens (box 0))
48 (define session-completion-tokens (box 0))
49
50 (define (mlx-reset-session-stats)
51 (call-with-semaphore stats-sema
52 (lambda ()
53 (set-box! session-prompt-tokens 0)
54 (set-box! session-completion-tokens 0))))
55
56 (define (mlx-accumulate-usage usage)
57 (when (and (hash? usage) (not (hash-empty? usage)))
58 (call-with-semaphore stats-sema
59 (lambda ()
60 (set-box! session-prompt-tokens
61 (+ (unbox session-prompt-tokens)
62 (hash-ref usage 'prompt_tokens 0)))
63 (set-box! session-completion-tokens
64 (+ (unbox session-completion-tokens)
65 (hash-ref usage 'completion_tokens 0)))))))
66
67 (define (mlx-print-session-stats)
68 (define-values (pt ct)
69 (call-with-semaphore stats-sema
70 (lambda ()
71 (values (unbox session-prompt-tokens)
72 (unbox session-completion-tokens)))))
73 (displayln "")
74 (displayln "Session token usage (local MLX -- no API cost):")
75 (displayln (format " Prompt tokens: ~a" pt))
76 (displayln (format " Completion tokens: ~a" ct))
77 (define mdl (or (provider-model (current-provider-json)) "?"))
78 (displayln (format " Estimated cost: $0 (local model ~a)" mdl)))
79
80 ;; ---------------------------------------------------------------------------
81 ;; Low-level POST (non-streaming, OpenAI-compatible)
82 ;;
83 ;; mlx_lm.server and Ollama's OpenAI shim both serve /v1/chat/completions.
84 ;; The request and response already match what chat-loop.rkt / chat-with-tools*
85 ;; consume, so we forward the OpenAI payload as-is and return the response
86 ;; unchanged. An optional Bearer key from the provider profile is sent when set
87 ;; (required by remote Ollama-style endpoints, ignored by a local server).
88
89 (define (mlx-api-key provider)
90 (define env-name (provider-api-key-env provider))
91 (and env-name
92 (getenv env-name)))
93
94 (define (post-mlx payload)
95 (define provider (current-provider-json))
96 ;; chat-loop.rkt already builds a complete OpenAI-shaped body, including
97 ;; tools/tool_choice when present, so it is forwarded as-is.
98 (define request* payload)
99 (define endpoint
100 (or (provider-endpoint provider)
101 (error 'mlx-serve
102 "active provider profile has no \"endpoint\"; set it in the harness config")))
103 (define key (mlx-api-key provider))
104 (define headers
105 (if (and key (not (string=? key "")))
106 (hash 'content-type "application/json"
107 'authorization (string-append "Bearer " key))
108 (hash 'content-type "application/json")))
109 (when (DEBUG-LOG)
110 (displayln (format "[DEBUG] mlx request (~a): ~a"
111 endpoint
112 (jsexpr->string request*))))
113 (define data
114 (with-handlers ([exn:fail? (lambda (e) (error 'mlx-serve "HTTP error: ~a" (exn-message e)))])
115 (define resp
116 (post endpoint
117 #:headers headers
118 #:json request*
119 #:timeouts (make-timeout-config #:request MLX-MAX-TIME
120 #:connect MLX-CONNECT-TIME)))
121 (define j (response-json resp))
122 (when (DEBUG-LOG)
123 (displayln (format "[DEBUG] mlx response: ~a" (jsexpr->string j))))
124 j))
125 (when (hash-has-key? data 'error)
126 (define err (hash-ref data 'error))
127 (define msg
128 (cond
129 [(hash? err) (hash-ref err 'message (format "~a" err))]
130 [else (format "~a" err)]))
131 (error 'mlx-serve "MLX API error: ~a" msg))
132 (unless (hash-has-key? data 'choices)
133 (error 'mlx-serve "MLX response has no 'choices'. Raw: ~a" (jsexpr->string data)))
134 (mlx-accumulate-usage (hash-ref data 'usage (hash)))
135 data)
136
137 ;; ---------------------------------------------------------------------------
138 ;; mlx-chat / mlx-chat-with-tools -- same signatures as fireworks-ai.rkt
139 ;;
140 ;; Generation defaults resolve through the profile in mlx-active-provider.
141
142 ;; Generation parameters resolve from the active profile; a missing model is an
143 ;; error, and missing generation parameters are simply left out of the request.
144
145 (define (m-gen-param key)
146 (generation-ref (provider-generation (current-provider-json)) key #f))
147
148 (define (m-model-id)
149 (or (provider-model (current-provider-json))
150 (error 'mlx-serve
151 "active provider profile has no \"model\"; set it in the harness config")))
152
153 (define (mlx-chat messages
154 #:model-id [model-id (m-model-id)]
155 #:max-tokens [max-tokens (m-gen-param 'max_tokens)]
156 #:temperature [temperature (m-gen-param 'temperature)])
157 (chat* post-mlx messages
158 #:model-id model-id
159 #:max-tokens max-tokens
160 #:temperature temperature))
161
162 (define (mlx-chat-with-tools messages tools
163 #:model-id [model-id (m-model-id)]
164 #:max-tokens [max-tokens (m-gen-param 'max_tokens)]
165 #:temperature [temperature (m-gen-param 'temperature)]
166 #:max-iterations [max-iterations 20])
167 (chat-with-tools* post-mlx messages tools
168 #:model-id model-id
169 #:max-tokens max-tokens
170 #:temperature temperature
171 #:max-iterations max-iterations))
No Conversion Functions Needed
Because mlx_lm.server speaks the OpenAI-compatible protocol end to end, mlx-serve.rkt needs no boundary conversion functions. Outgoing messages and incoming responses already carry tool-call arguments as JSON strings and token counts as prompt_tokens / completion_tokens, exactly the shape the shared loop expects.
The module keeps one parameter, mlx-active-provider, which holds the provider profile that requests should read. agent.rkt binds it with parameterize around each call, so the endpoint, the model, and the generation settings all resolve from the active profile instead of from module-level constants. A profile with no endpoint or no model is an error rather than a silent fallback to a hard-coded default. An optional Bearer key from the profile’s api_key_env is sent when one is set, which is what a remote Ollama-style endpoint needs; a purely local server ignores it.
Because local inference is free, the cost display is always $0. The stats are informational only.
Hierarchical Provider Configuration
Earlier versions of this harness selected a provider with a single environment variable. This version moves every provider-specific value into configuration, in the rough style of the Pi coding harness. Two JSON layers are merged at startup, and the project-local file wins over the global one:
- Global:
~/.coding_harness.json, the base configuration. - Local:
.local_coding_harness.jsonin the current directory, an optional per-project override.
A minimal config declares one or more named providers and a default:
1 {
2 "default_provider": "mlx",
3 "providers": {
4 "mlx": {
5 "type": "mlx",
6 "endpoint": "http://localhost:11434/v1/chat/completions",
7 "model": "mlx-community/gemma-4-26B-A4B-it-OptiQ-4bit",
8 "generation": { "temperature": 0.6, "max_tokens": 32768 }
9 },
10 "fireworks": {
11 "type": "openai",
12 "endpoint": "https://api.fireworks.ai/inference/v1/chat/completions",
13 "api_key_env": "FIREWORKS_API_KEY",
14 "model": "accounts/fireworks/models/deepseek-v4p1-flash",
15 "generation": { "temperature": 0.6, "max_tokens": 32768 },
16 "pricing": { "input": 0.14, "cached_input": 0.028, "output": 0.28 }
17 }
18 }
19 }
The type field selects the wire format: "mlx" for the local OpenAI-compatible server (the strings "ollama", "omlx", and "sushi" are accepted as aliases and mapped to the same backend) and "openai" for Fireworks and any compatible endpoint. The optional pricing block supplies the per-million-token rates that /tokens uses. Nothing in the Racket code names a provider, an endpoint, or a price.
Here is the complete module:
1 #lang racket
2
3 (require json
4 racket/file
5 racket/string)
6
7 (provide load-harness-config
8 harness-config
9 config-provider-names
10 config-provider
11 config-active-provider-name
12 config-set-active-provider!
13 config-active-provider
14 provider-type
15 provider-endpoint
16 provider-model
17 provider-api-key-env
18 provider-generation
19 generation-ref
20 provider-pricing
21 pricing-ref
22 print-config-summary)
23
24 ;; ---------------------------------------------------------------------------
25 ;; JSON loading helpers
26
27 (define (read-json-file path)
28 ;; -> jsexpr hash, or #f if the file is missing / unreadable / not an object
29 (with-handlers ([exn:fail? (lambda (_) #f)])
30 (and (file-exists? path)
31 (let ([v (with-input-from-file path read-json)])
32 (and (hash? v) v)))))
33
34 (define (global-config-path)
35 (build-path (find-system-path 'home-dir) ".coding_harness.json"))
36
37 (define (local-config-path)
38 (build-path (current-directory) ".local_coding_harness.json"))
39
40 ;; ---------------------------------------------------------------------------
41 ;; Recursive hash merge (local overrides global)
42
43 (define (deep-merge global local)
44 (cond
45 [(not (hash? global)) local]
46 [(not (hash? local)) local]
47 [else
48 (for/fold ([acc global])
49 ([(k v) (in-hash local)])
50 (hash-set acc k
51 (if (and (hash? (hash-ref acc k #f)) (hash? v))
52 (deep-merge (hash-ref acc k) v)
53 v)))]))
54
55 ;; ---------------------------------------------------------------------------
56 ;; The merged config, loaded once at startup (reloadable via load-harness-config)
57
58 (define harness-config (make-parameter (hash)))
59
60 (define (load-harness-config)
61 ;; Load global then local, deep-merge, store in the parameter, and return it.
62 (define global (or (read-json-file (global-config-path)) (hash)))
63 (define local (or (read-json-file (local-config-path)) (hash)))
64 (define merged (deep-merge global local))
65 (harness-config merged)
66 merged)
67
68 ;; ---------------------------------------------------------------------------
69 ;; Providers
70
71 (define (config-providers)
72 (define p (hash-ref (harness-config) 'providers (hash)))
73 (if (hash? p) p (hash)))
74
75 (define (config-provider-names)
76 (sort (map symbol->string (hash-keys (config-providers))) string<?))
77
78 (define (name->key name)
79 ;; provider sections are keyed by the profile name; JSON object keys come
80 ;; back as symbols, so accept either a string or symbol name.
81 (cond
82 [(symbol? name) name]
83 [(string? name) (string->symbol name)]
84 [else (string->symbol (format "~a" name))]))
85
86 (define (config-provider name)
87 ;; -> provider hash for profile `name`, or #f
88 (hash-ref (config-providers) (name->key name) #f))
89
90 ;; Active provider profile ---------------------------------------------------
91
92 ;; Mutable cell: the name of the provider section currently in use. Defaults
93 ;; to "default_provider" from config, else "fireworks" if that section exists,
94 ;; else the first declared provider, else #f (use the compiled-in defaults of
95 ;; fireworks-ai.rkt / mlx-serve.rkt).
96 (define active-provider-name (box #f))
97
98 (define (pick-default-provider-name cfg)
99 (define declared (hash-ref cfg 'default_provider #f))
100 (define names (map symbol->string (hash-keys (config-providers))))
101 (cond
102 [(and (string? declared)
103 (hash-has-key? (config-providers) (string->symbol declared)))
104 declared]
105 [(member "fireworks" names) "fireworks"]
106 [(pair? names) (first (sort names string<?))]
107 [else #f]))
108
109 (define (config-active-provider-name)
110 (or (unbox active-provider-name)
111 (let ([n (pick-default-provider-name (harness-config))])
112 (set-box! active-provider-name n)
113 n)))
114
115 (define (config-set-active-provider! name)
116 (when (and name (config-provider name))
117 (set-box! active-provider-name
118 (if (string? name) name (symbol->string name))))
119 (unbox active-provider-name))
120
121 (define (config-active-provider)
122 ;; -> provider hash of the active profile, or #f when there is no config
123 (define n (config-active-provider-name))
124 (and n (config-provider n)))
125
126 ;; ---------------------------------------------------------------------------
127 ;; Provider field accessors (all tolerant of missing keys)
128
129 (define (provider-type provider)
130 ;; -> 'mlx | 'openai -- defaults to 'openai
131 ;; "mlx" selects the local mlx-serve backend (formerly "ollama"); "ollama",
132 ;; "omlx", and "sushi" are also accepted and mapped to 'mlx for compatibility.
133 (define t (and provider (hash-ref provider 'type #f)))
134 (define low
135 (and (or (string? t) (symbol? t))
136 (string-downcase (format "~a" t))))
137 (cond
138 [(member low '("mlx" "ollama" "omlx" "sushi")) 'mlx]
139 [else 'openai]))
140
141 (define (provider-endpoint provider)
142 (define e (and provider (hash-ref provider 'endpoint #f)))
143 (and (string? e) (not (string=? e "")) e))
144
145 (define (provider-model provider)
146 (define m (and provider (hash-ref provider 'model #f)))
147 (and (string? m) (not (string=? m "")) m))
148
149 (define (provider-api-key-env provider)
150 ;; Name of the env var holding the Bearer key for this endpoint, or #f.
151 ;; Absent/empty/null means "no key" (plain local MLX).
152 (define k (and provider (hash-ref provider 'api_key_env #f)))
153 (and (string? k) (not (string=? k "")) k))
154
155 (define (provider-generation provider)
156 (define g (and provider (hash-ref provider 'generation #f)))
157 (if (hash? g) g (hash)))
158
159 (define (provider-pricing provider)
160 ;; -> hash of per-1M-token USD rates ('input, 'cached_input, 'output), or an
161 ;; empty hash when the profile declares none. Rates live in the provider
162 ;; profile so that none are compiled into the code.
163 (define g (and provider (hash-ref provider 'pricing #f)))
164 (if (hash? g) g (hash)))
165
166 (define (pricing-ref pricing key)
167 ;; -> number, or #f when the profile does not declare that rate. The #f
168 ;; result means "unknown", which callers report rather than guessing a value.
169 (cond
170 [(not (hash? pricing)) #f]
171 [(hash-has-key? pricing key) (hash-ref pricing key)]
172 [(hash-has-key? pricing (string->symbol (format "~a" key)))
173 (hash-ref pricing (string->symbol (format "~a" key)))]
174 [(hash-has-key? pricing (format "~a" key)) (hash-ref pricing (format "~a" key))]
175 [else #f]))
176
177 (define (generation-ref generation key default)
178 ;; Fetch a generation parameter ("temperature", "max_tokens", "think", ...)
179 ;; accepting symbol or string keys because JSON may give either.
180 (cond
181 [(not (hash? generation)) default]
182 [(hash-has-key? generation key) (hash-ref generation key)]
183 [(hash-has-key? generation (string->symbol (format "~a" key)))
184 (hash-ref generation (string->symbol (format "~a" key)))]
185 [(hash-has-key? generation (format "~a" key))
186 (hash-ref generation (format "~a" key))]
187 [else default]))
188
189 ;; ---------------------------------------------------------------------------
190 ;; Debug helper
191
192 (define (print-config-summary)
193 (define cfg (harness-config))
194 (displayln (format "Config files: ~a ~a / ~a ~a"
195 (global-config-path)
196 (if (file-exists? (global-config-path)) "(loaded)" "(absent)")
197 (local-config-path)
198 (if (file-exists? (local-config-path)) "(loaded)" "(absent)")))
199 (displayln (format "Providers: ~a"
200 (string-join (config-provider-names) ", ")))
201 (displayln (format "Active: ~a" (or (config-active-provider-name) "(defaults)"))))
Merging the Two Layers
load-harness-config reads both files, deep-merges them, and stores the result in the harness-config parameter. deep-merge recurses into nested hashes so a local file can override a single field, such as a model id, without restating the whole provider; anything that is not a hash, such as a string, a number, or a list, is replaced wholesale by the local value. A missing or malformed file is treated as an empty hash, so the harness starts with whatever configuration is actually present.
The active provider is a name kept in a box. On first use, config-active-provider-name picks default_provider if it names a real profile, otherwise fireworks if that profile exists, otherwise the alphabetically first profile. /provider, --provider, and the environment can all change it at run time through config-set-active-provider!.
The accessors near the bottom of the file, provider-endpoint, provider-model, provider-api-key-env, provider-generation, and provider-pricing, are deliberately tolerant: each returns #f or an empty hash when the field is absent, and the caller decides whether that is fatal. The two clients treat a missing endpoint or model as an error, so a half-written profile produces a clear message instead of a request sent to #f. pricing-ref and generation-ref accept either a symbol or a string key, because read-json returns JSON object keys as symbols while a hand-written config might use strings.
The Tool Registry
Defining and Rendering Tools
tools.rkt maintains a central hash table of all registered tools. Here is the complete file:
1 #lang racket
2
3 (require racket/file
4 racket/port
5 racket/string
6 racket/system
7 racket/list
8 json
9 "approval.rkt")
10
11 (provide define-tool
12 render-tools
13 execute-tool-calls
14 register-all
15 ENABLED-TOOLS
16 auto-approve?
17 dry-run?
18 quiet-mode?)
19
20 ;; ---------------------------------------------------------------------------
21 ;; Registry
22
23 (define registry (make-hash))
24
25 (define SHELL-WHITELIST (set "make" "ls" "pwd" "cat" "uv"))
26 (define MAX-CHECK-OUTPUT-CHARS 2000)
27
28 ;; CLI-controlled modes
29 (define auto-approve? (make-parameter #f))
30 (define dry-run? (make-parameter #f))
31 (define quiet-mode? (make-parameter #f))
32
33 (define (define-tool name params description handler)
34 ;; params : list of (list pname ptype pdesc)
35 (hash-set! registry name
36 (hash 'name name
37 'description description
38 'parameters params
39 'handler handler)))
40
41 (define (render-tools names)
42 (for/list ([name (in-list names)])
43 (define tool (hash-ref registry name #f))
44 (unless tool (error 'render-tools "Undefined tool: ~a" name))
45 (define props (make-hash))
46 (define required '())
47 (for ([p (in-list (hash-ref tool 'parameters))])
48 (define pname (first p))
49 (define ptype (second p))
50 (define pdesc (third p))
51 (hash-set! props (string->symbol pname)
52 (hash 'type ptype 'description pdesc))
53 ;; required must be a JSON array of strings, not symbols
54 (set! required (cons pname required)))
55 (hash 'type "function"
56 'function (hash 'name (hash-ref tool 'name)
57 'description (hash-ref tool 'description)
58 'parameters (hash 'type "object"
59 'properties props
60 'required (reverse required))))))
61
62 ;; ---------------------------------------------------------------------------
63 ;; Tool dispatch
64
65 (define (call-tool name args)
66 (define tool (hash-ref registry name #f))
67 (unless tool (error 'call-tool "Unknown tool: ~a" name))
68 (define params (hash-ref tool 'parameters))
69 ;; Missing required args? Return an actionable error describing the expected
70 ;; argument list -- small models frequently emit malformed/truncated
71 ;; arguments, and silently receiving #f tends to send them into retry loops.
72 (define missing
73 (for/list ([p (in-list params)]
74 #:when (not (hash-ref args (string->symbol (first p)) #f)))
75 (first p)))
76 (cond
77 [(pair? missing)
78 (format "Error: tool '~a' missing required argument(s): ~a. Expected arguments (JSON object): ~a"
79 name
80 (string-join missing ", ")
81 (string-join (for/list ([p (in-list params)]) (first p)) ", "))]
82 [else
83 (define positional
84 (for/list ([p (in-list params)])
85 (hash-ref args (string->symbol (first p)) #f)))
86 (with-handlers ([exn:fail? (lambda (e) (format "Error: tool '~a' raised: ~a (check argument types/values)"
87 name (exn-message e)))])
88 (define result (apply (hash-ref tool 'handler) positional))
89 (if result (format "~a" result) ""))]))
90
91 (define (execute-tool-calls tool-calls)
92 ;; tool-calls : list of hashes with 'id, 'function {name, arguments}
93 ;; Returns list of (list call-id name result-str)
94 (define results '())
95 (for ([call (in-list tool-calls)])
96 (define call-id (hash-ref call 'id ""))
97 (define func (hash-ref call 'function (hash)))
98 (define name (hash-ref func 'name ""))
99 (define args-json (hash-ref func 'arguments "{}"))
100 (define short
101 (if (<= (string-length args-json) 120)
102 args-json
103 (string-append (substring args-json 0 117) "...")))
104 (unless (quiet-mode?)
105 (displayln (format "* ~a ~a" name short)))
106 (define args-parsed
107 (with-handlers ([exn:fail? (lambda (_) 'BAD-JSON)])
108 (let ([j (string->jsexpr args-json)])
109 (if (hash? j) j 'NOT-OBJECT))))
110 (define result
111 (cond
112 ;; Truncated tool call -- the model stopped mid-generation, so no
113 ;; function name survived. Feed that back instead of crashing.
114 [(string=? (string-trim name) "")
115 (format "Error: the model's tool call was truncated mid-generation (no function name provided). Received arguments: ~a"
116 short)]
117 [(eq? args-parsed 'BAD-JSON)
118 (format "Error: invalid JSON in arguments for tool '~a'. Received: ~a"
119 name short)]
120 [(eq? args-parsed 'NOT-OBJECT)
121 (format "Error: arguments for tool '~a' must be a JSON object. Received: ~a"
122 name short)]
123 [else
124 ;; Unknown tool names, contract violations, etc. become feedback to the
125 ;; model rather than an uncaught exception that aborts the loop.
126 (with-handlers ([exn:fail? (lambda (e)
127 (format "Error: tool '~a' raised: ~a"
128 name (exn-message e)))])
129 (call-tool name args-parsed))]))
130 (set! results (append results (list (list call-id name result)))))
131 results)
132
133 ;; ---------------------------------------------------------------------------
134 ;; Helpers: run subprocess and capture combined output
135
136 (define (run-external exe args)
137 ;; exe : string, args : (listof string) -> (values combined-output exit-code)
138 (define exe-path (find-executable-path exe))
139 (unless exe-path
140 (error 'run-external "Executable not found: ~a" exe))
141 (define-values (sp stdout stdin stderr)
142 (apply subprocess #f #f #f exe-path args))
143 (close-output-port stdin)
144 (define out-str (port->string stdout))
145 (define err-str (port->string stderr))
146 (close-input-port stdout)
147 (close-input-port stderr)
148 (subprocess-wait sp)
149 (define status (subprocess-status sp))
150 (define code (if (number? status) status 1))
151 (values (string-append out-str err-str) code))
152
153 (define (shell-quote s)
154 (string-append "'" (string-replace s "'" "'\\''") "'"))
155
156 (define (truncate-string s max-len)
157 (if (> (string-length s) max-len)
158 (string-append (substring s 0 max-len)
159 (format "\n... (truncated, ~a total chars)" (string-length s)))
160 s))
161
162 ;; Hidden files (ignored from listings and reject read attempts):
163 ;; - names ending in ~ (e.g. foo.rkt~)
164 ;; - names wrapped in #...# (e.g. #foo.rkt#)
165 ;; - names starting with . (e.g. .git, .gitignore, .env)
166 (define (hidden-file? name)
167 (or (string-suffix? name "~")
168 (and (string-prefix? name "#")
169 (string-suffix? name "#"))
170 (string-prefix? name ".")))
171
172 ;; ---------------------------------------------------------------------------
173 ;; Tool implementations
174
175 (define (tool-read-file path)
176 (with-handlers ([exn:fail? (lambda (e) (format "Error reading ~a: ~a" path (exn-message e)))])
177 (define fname (path->string (file-name-from-path path)))
178 (if (hidden-file? fname)
179 (format "refusing to read hidden/internal file: ~a" path)
180 (file->string path))))
181
182 (define (tool-list-dir path)
183 (with-handlers ([exn:fail? (lambda (e) (format "Error listing ~a: ~a" path (exn-message e)))])
184 (define entries (directory-list path))
185 (define lines
186 (for/list ([e (in-list (sort (map path->string entries) string<?))]
187 #:unless (hidden-file? (path->string e)))
188 (define full (build-path path e))
189 (if (directory-exists? full)
190 (string-append e "/")
191 e)))
192 (string-join lines "\n")))
193
194 (define (tool-grep pattern path)
195 (with-handlers ([exn:fail? (lambda (e) (format "Error running grep: ~a" (exn-message e)))])
196 (define-values (out code) (run-external "grep" (list "-rnE" pattern path)))
197 out))
198
199 (define (strip-shell-quotes s)
200 (if (and (>= (string-length s) 2)
201 (let ([first (string-ref s 0)]
202 [last (string-ref s (sub1 (string-length s)))])
203 (or (and (char=? first #\") (char=? last #\"))
204 (and (char=? first #\') (char=? last #\')))))
205 (substring s 1 (sub1 (string-length s)))
206 s))
207
208 (define (hidden-arg? s)
209 (define cleaned (strip-shell-quotes s))
210 (and (not (string-prefix? cleaned "-"))
211 (hidden-file? (path->string (file-name-from-path cleaned)))))
212
213 (define (filter-ls-output out)
214 (define lines (string-split out "\n"))
215 (define filtered
216 (for/list ([line (in-list lines)]
217 #:unless (let ([t (string-trim line)])
218 (or (string-prefix? t "total ")
219 (equal? t ""))))
220 (define trimmed (string-trim line))
221 (define tokens (string-split trimmed))
222 (cond
223 [(null? tokens) #f]
224 [(regexp-match? #rx"^[-d]" (first tokens))
225 (define fname (last tokens))
226 (if (or (hidden-file? fname) (member fname '("." "..")))
227 #f
228 line)]
229 [else
230 (if (hidden-file? trimmed) #f line)])))
231 (define result-lines (for/list ([f (in-list filtered)] #:when f) f))
232 (if (null? result-lines) "" (string-join result-lines "\n")))
233
234 (define (tool-run-shell command)
235 (define tokens (string-split (string-trim command)))
236 (cond
237 [(null? tokens) "empty command"]
238 [else
239 (define cmd (first tokens))
240 (cond
241 [(not (set-member? SHELL-WHITELIST cmd))
242 (format "Command '~a' not whitelisted. Allowed: ~a"
243 cmd (string-join (sort (set->list SHELL-WHITELIST) string<?) ", "))]
244 [(and (not (equal? cmd "ls"))
245 (ormap hidden-arg? (rest tokens)))
246 => (lambda (bad)
247 (format "refusing to run command referencing hidden/internal file: ~a" bad))]
248 [else
249 (with-handlers ([exn:fail? (lambda (e) (format "Error running command: ~a" (exn-message e)))])
250 (define args (rest tokens))
251 (define-values (out code) (run-external cmd args))
252 (define filtered-out (if (equal? cmd "ls") (filter-ls-output out) out))
253 (string-append filtered-out (format "(exit ~a)" code)))])]))
254
255 (define (run-make-check)
256 (with-handlers ([exn:fail? (lambda (e) (values (format "make check error: ~a" (exn-message e)) 1))])
257 (run-external "make" (list "check"))))
258
259 (define (tool-propose-edit path old new)
260 (define exists? (file-exists? path))
261 (define current
262 (if exists?
263 (with-handlers ([exn:fail? (lambda (e) (format "Error reading ~a: ~a" path (exn-message e)))])
264 (file->string path))
265 ""))
266 ;; If read failed and returned error string, treat as error
267 (when (and exists? (string-prefix? current "Error reading"))
268 current)
269 (cond
270 [(and exists? (not (string=? current old)))
271 (format "stale base: on-disk contents of ~a do not match the 'old' you provided. Read the file again and retry." path)]
272 [(and exists? (string=? current new))
273 "no changes (proposed content matches current file)"]
274 [(and (not exists?) (string=? new ""))
275 "refused: cannot create an empty file"]
276 [else
277 (define diff-text (unified-diff current new (string-append "a/" path) (string-append "b/" path)))
278 (displayln "")
279 (unless exists? (displayln (format "(new file: ~a)" path)))
280 (print-colored-diff diff-text)
281 (cond
282 [(dry-run?)
283 "dry-run: diff shown, file not written (use without --dry-run to apply)"]
284 [(auto-approve?)
285 ;; Safety: still show diff above, then auto-apply without prompting
286 (unless (quiet-mode?)
287 (displayln "[auto-approve: applying change without prompt]"))
288 (make-parent-directory* path)
289 (call-with-output-file path #:exists 'truncate
290 (lambda (out) (display new out)))
291 (define-values (out status) (run-make-check))
292 (if (= status 0)
293 "applied (auto-approved); make check passed"
294 (format "applied (auto-approved); make check FAILED (exit ~a):\n~a"
295 status (truncate-string out MAX-CHECK-OUTPUT-CHARS)))]
296 [else
297 (define answer (prompt-yes-no-skip))
298 (cond
299 [(eq? answer 'no) "user rejected the change"]
300 [(eq? answer 'skip)
301 (define reason (prompt-reason))
302 (format "user skipped: ~a" reason)]
303 [else ; 'yes
304 (make-parent-directory* path)
305 (call-with-output-file path #:exists 'truncate
306 (lambda (out) (display new out)))
307 (define-values (out status) (run-make-check))
308 (if (= status 0)
309 "applied; make check passed"
310 (format "applied; make check FAILED (exit ~a):\n~a"
311 status (truncate-string out MAX-CHECK-OUTPUT-CHARS)))])])]))
312
313 ;; ---------------------------------------------------------------------------
314 ;; Registration
315
316 (define (register-all)
317 (define-tool
318 "read_file"
319 (list (list "path" "string" "File path relative to the working directory."))
320 "Read and return the contents of a file. Refuses to read hidden/internal files (~, #...#, and dotfiles)."
321 tool-read-file)
322 (define-tool
323 "list_dir"
324 (list (list "path" "string" "Directory path. Use \".\" for the working directory."))
325 "List files and subdirectories (with trailing /) in a directory. Hidden/internal files (~, #...#, and dotfiles) are excluded."
326 tool-list-dir)
327 (define-tool
328 "grep"
329 (list (list "pattern" "string" "Extended regex pattern to search for.")
330 (list "path" "string" "Directory or file path to search."))
331 "Recursively grep files for PATTERN. Wraps `grep -rnE`."
332 tool-grep)
333 (define-tool
334 "run_shell"
335 (list (list "command" "string" "Shell command. Only whitelisted commands may run: make, ls, pwd, cat, uv."))
336 "Run a whitelisted shell command and return its combined output. Refuses commands that reference hidden/internal files."
337 tool-run-shell)
338 (define-tool
339 "propose_edit"
340 (list (list "path" "string" "Path to the file to edit or create.")
341 (list "old" "string" "For an existing file: the exact current contents. For a new file: pass empty string.")
342 (list "new" "string" "The proposed new contents of the file, in full."))
343 "Propose an edit or new-file creation. The user is shown a unified diff and asked to approve. On approval the file is written and `make check` is run."
344 tool-propose-edit))
345
346 (define ENABLED-TOOLS (list "read_file" "list_dir" "grep" "run_shell" "propose_edit"))
Each tool is stored as a hash with its name, description, parameter list, and handler function. render-tools converts the registry entries into OpenAI function-calling schema format: a list of hashes the API understands as callable functions. The model receives these alongside the conversation and decides which, if any, to invoke.
Tool Dispatch
execute-tool-calls receives the list of tool call objects from the API response and dispatches each one:
1 (define (execute-tool-calls tool-calls)
2 ;; tool-calls : list of hashes with 'id, 'function {name, arguments}
3 ;; Returns list of (list call-id name result-str)
4 (define results '())
5 (for ([call (in-list tool-calls)])
6 (define call-id (hash-ref call 'id ""))
7 (define func (hash-ref call 'function (hash)))
8 (define name (hash-ref func 'name ""))
9 (define args-json (hash-ref func 'arguments "{}"))
10 (define short
11 (if (<= (string-length args-json) 120)
12 args-json
13 (string-append (substring args-json 0 117) "...")))
14 (unless (quiet-mode?)
15 (displayln (format "* ~a ~a" name short)))
16 (define args-parsed
17 (with-handlers ([exn:fail? (lambda (_) 'BAD-JSON)])
18 (let ([j (string->jsexpr args-json)])
19 (if (hash? j) j 'NOT-OBJECT))))
20 (define result
21 (cond
22 ;; Truncated tool call -- the model stopped mid-generation, so no
23 ;; function name survived. Feed that back instead of crashing.
24 [(string=? (string-trim name) "")
25 (format "Error: the model's tool call was truncated mid-generation (no function name provided). Received arguments: ~a"
26 short)]
27 [(eq? args-parsed 'BAD-JSON)
28 (format "Error: invalid JSON in arguments for tool '~a'. Received: ~a"
29 name short)]
30 [(eq? args-parsed 'NOT-OBJECT)
31 (format "Error: arguments for tool '~a' must be a JSON object. Received: ~a"
32 name short)]
33 [else
34 ;; Unknown tool names, contract violations, etc. become feedback to the
35 ;; model rather than an uncaught exception that aborts the loop.
36 (with-handlers ([exn:fail? (lambda (e)
37 (format "Error: tool '~a' raised: ~a"
38 name (exn-message e)))])
39 (call-tool name args-parsed))]))
40 (set! results (append results (list (list call-id name result)))))
41 results)
Each call prints its name and a truncated copy of its arguments unless quiet mode is on, so the user can watch what the model is doing. These error branches matter more than they look. A weak model can truncate a tool call mid-generation, which leaves the function name empty; it can emit arguments that are not valid JSON; or it can emit valid JSON that is not an object. Each of those cases becomes an explanatory string that goes back to the model as the tool result. A raised exception inside a handler is caught the same way. The loop therefore keeps running and the model gets a chance to correct itself, rather than the whole session dying on one malformed call.
The Five Coding Tools
The agent registers five tools at startup:
| Tool | Purpose |
|---|---|
read_file |
Return the full text of a file |
list_dir |
List files and subdirectories in a directory |
grep |
Recursively search for an extended regex pattern |
run_shell |
Run a whitelisted shell command and return its output |
propose_edit |
Show a colored diff and ask the user to approve the change |
Do not let the short list suggest that the model can wander the filesystem. read_file, list_dir, and run_shell share a hidden-file? predicate that treats editor backups ending in ~, Emacs lock files wrapped in #...#, and dotfiles as off limits. read_file refuses them, list_dir filters them out, and run_shell both refuses any argument that names one and filters ls output. run_shell also enforces a strict command whitelist so the model cannot run arbitrary shell commands:
1 (define (hidden-arg? s)
2 (define cleaned (strip-shell-quotes s))
3 (and (not (string-prefix? cleaned "-"))
4 (hidden-file? (path->string (file-name-from-path cleaned)))))
5
6 (define (filter-ls-output out)
7 (define lines (string-split out "\n"))
8 (define filtered
9 (for/list ([line (in-list lines)]
10 #:unless (let ([t (string-trim line)])
11 (or (string-prefix? t "total ")
12 (equal? t ""))))
13 (define trimmed (string-trim line))
14 (define tokens (string-split trimmed))
15 (cond
16 [(null? tokens) #f]
17 [(regexp-match? #rx"^[-d]" (first tokens))
18 (define fname (last tokens))
19 (if (or (hidden-file? fname) (member fname '("." "..")))
20 #f
21 line)]
22 [else
23 (if (hidden-file? trimmed) #f line)])))
24 (define result-lines (for/list ([f (in-list filtered)] #:when f) f))
25 (if (null? result-lines) "" (string-join result-lines "\n")))
26
27 (define (tool-run-shell command)
28 (define tokens (string-split (string-trim command)))
29 (cond
30 [(null? tokens) "empty command"]
31 [else
32 (define cmd (first tokens))
33 (cond
34 [(not (set-member? SHELL-WHITELIST cmd))
35 (format "Command '~a' not whitelisted. Allowed: ~a"
36 cmd (string-join (sort (set->list SHELL-WHITELIST) string<?) ", "))]
37 [(and (not (equal? cmd "ls"))
38 (ormap hidden-arg? (rest tokens)))
39 => (lambda (bad)
40 (format "refusing to run command referencing hidden/internal file: ~a" bad))]
41 [else
42 (with-handlers ([exn:fail? (lambda (e) (format "Error running command: ~a" (exn-message e)))])
43 (define args (rest tokens))
44 (define-values (out code) (run-external cmd args))
45 (define filtered-out (if (equal? cmd "ls") (filter-ls-output out) out))
46 (string-append filtered-out (format "(exit ~a)" code)))])]))
If the model attempts a disallowed command, or names a file the agent considers internal, it receives an error string describing what is allowed. It can then adapt its approach rather than causing the agent to crash. The whitelist currently allows make, ls, pwd, cat, and uv (the Python package runner). Everything else is refused.
The propose_edit Approval Gate
propose_edit is the most critical tool. Before writing any file it checks for several error conditions, shows the user a diff, waits for approval or applies the change automatically, and then runs make check:
1 (define (tool-propose-edit path old new)
2 (define exists? (file-exists? path))
3 (define current
4 (if exists?
5 (with-handlers ([exn:fail? (lambda (e) (format "Error reading ~a: ~a" path (exn-message e)))])
6 (file->string path))
7 ""))
8 ;; If read failed and returned error string, treat as error
9 (when (and exists? (string-prefix? current "Error reading"))
10 current)
11 (cond
12 [(and exists? (not (string=? current old)))
13 (format "stale base: on-disk contents of ~a do not match the 'old' you provided. Read the file again and retry." path)]
14 [(and exists? (string=? current new))
15 "no changes (proposed content matches current file)"]
16 [(and (not exists?) (string=? new ""))
17 "refused: cannot create an empty file"]
18 [else
19 (define diff-text (unified-diff current new (string-append "a/" path) (string-append "b/" path)))
20 (displayln "")
21 (unless exists? (displayln (format "(new file: ~a)" path)))
22 (print-colored-diff diff-text)
23 (cond
24 [(dry-run?)
25 "dry-run: diff shown, file not written (use without --dry-run to apply)"]
26 [(auto-approve?)
27 ;; Safety: still show diff above, then auto-apply without prompting
28 (unless (quiet-mode?)
29 (displayln "[auto-approve: applying change without prompt]"))
30 (make-parent-directory* path)
31 (call-with-output-file path #:exists 'truncate
32 (lambda (out) (display new out)))
33 (define-values (out status) (run-make-check))
34 (if (= status 0)
35 "applied (auto-approved); make check passed"
36 (format "applied (auto-approved); make check FAILED (exit ~a):\n~a"
37 status (truncate-string out MAX-CHECK-OUTPUT-CHARS)))]
38 [else
39 (define answer (prompt-yes-no-skip))
40 (cond
41 [(eq? answer 'no) "user rejected the change"]
42 [(eq? answer 'skip)
43 (define reason (prompt-reason))
44 (format "user skipped: ~a" reason)]
45 [else ; 'yes
46 (make-parent-directory* path)
47 (call-with-output-file path #:exists 'truncate
48 (lambda (out) (display new out)))
49 (define-values (out status) (run-make-check))
50 (if (= status 0)
51 "applied; make check passed"
52 (format "applied; make check FAILED (exit ~a):\n~a"
53 status (truncate-string out MAX-CHECK-OUTPUT-CHARS)))])])]))
The stale-base guard is worth understanding carefully. The model reads a file, then constructs a proposed edit based on that content. If the user edits the file externally in between, a naive tool would overwrite those changes silently. By requiring old to exactly match what is on disk, the tool forces the model to re-read the file before retrying, and the mismatch is reported as a tool result the model can read and respond to. The neighboring guards reject a no-op edit whose new content already matches the file, and refuse to create an empty file.
The three application paths share the same write and check code. --dry-run stops after the diff. --yes prints the diff and applies it without prompting. The default path asks the user for y, n, or s, and on s it collects a one-line reason so the model learns why the change was skipped.
The make check gate closes another important feedback loop. If the edit compiles cleanly, "applied; make check passed" goes back into the conversation history and the model can proceed. If make check fails, the output goes back as well, giving the model the compiler errors it needs to self-correct on the next turn. The output is truncated to MAX-CHECK-OUTPUT-CHARS characters so a huge build log does not blow up the context window.
The Approval and Diff System
Generating a Unified Diff
approval.rkt generates diffs by writing the two file versions to temporary files and calling the system diff -u utility. Here is the full source file:
1 #lang racket
2
3 (require racket/file
4 racket/port
5 racket/string
6 racket/system)
7
8 (provide unified-diff
9 print-colored-diff
10 prompt-yes-no-skip
11 prompt-reason
12 color-enabled?)
13
14 ;; ---------------------------------------------------------------------------
15 ;; ANSI colours
16
17 (define ANSI-RED "\033[31m")
18 (define ANSI-GREEN "\033[32m")
19 (define ANSI-CYAN "\033[36m")
20 (define ANSI-RESET "\033[0m")
21
22 ;; When #f, print diffs without ANSI (for --plain / --no-color / piped output)
23 (define color-enabled? (make-parameter #t))
24
25 ;; ---------------------------------------------------------------------------
26 ;; Shell quoting helper (single-quote, escape embedded single quotes)
27
28 (define (shell-quote s)
29 (string-append "'"
30 (string-replace s "'" "'\\''")
31 "'"))
32
33 ;; ---------------------------------------------------------------------------
34 ;; unified-diff : string string string string -> string
35 ;; Runs `diff -u` on two temporary files and returns stdout.
36
37 (define (unified-diff old-content new-content old-label new-label)
38 (define old-path (make-temporary-file "rk-diff-old~a"))
39 (define new-path (make-temporary-file "rk-diff-new~a"))
40 (define out-path (make-temporary-file "rk-diff-out~a"))
41 (dynamic-wind
42 void
43 (lambda ()
44 (call-with-output-file old-path #:exists 'truncate
45 (lambda (out) (display old-content out)))
46 (call-with-output-file new-path #:exists 'truncate
47 (lambda (out) (display new-content out)))
48 (define cmd
49 (format "diff -u -L ~a -L ~a ~a ~a > ~a 2>&1"
50 (shell-quote old-label)
51 (shell-quote new-label)
52 (shell-quote (path->string old-path))
53 (shell-quote (path->string new-path))
54 (shell-quote (path->string out-path))))
55 (system cmd)
56 (with-handlers ([exn:fail? (lambda (_) "")])
57 (file->string out-path)))
58 (lambda ()
59 (when (file-exists? old-path) (delete-file old-path))
60 (when (file-exists? new-path) (delete-file new-path))
61 (when (file-exists? out-path) (delete-file out-path)))))
62
63 ;; ---------------------------------------------------------------------------
64 ;; print-colored-diff : string -> void
65
66 (define (print-colored-diff diff-text)
67 (for ([line (in-list (string-split diff-text "\n"))])
68 (cond
69 [(color-enabled?)
70 (cond
71 [(or (string-prefix? line "+++")
72 (string-prefix? line "---")
73 (string-prefix? line "@@"))
74 (displayln (string-append ANSI-CYAN line ANSI-RESET))]
75 [(string-prefix? line "+")
76 (displayln (string-append ANSI-GREEN line ANSI-RESET))]
77 [(string-prefix? line "-")
78 (displayln (string-append ANSI-RED line ANSI-RESET))]
79 [else (displayln line)])]
80 [else (displayln line)])))
81
82 ;; ---------------------------------------------------------------------------
83 ;; prompt-yes-no-skip : -> (or 'yes 'no 'skip)
84
85 (define (prompt-yes-no-skip)
86 (let loop ()
87 (display "\nApply this change? [y]es / [n]o / [s]kip and tell the model why: ")
88 (flush-output)
89 (define line (read-line (current-input-port)))
90 (define norm (if (eof-object? line) "" (string-downcase (string-trim line))))
91 (cond
92 [(member norm '("y" "yes")) 'yes]
93 [(member norm '("n" "no")) 'no]
94 [(member norm '("s" "skip")) 'skip]
95 [else
96 (displayln "Please answer y, n, or s.")
97 (loop)])))
98
99 ;; ---------------------------------------------------------------------------
100 ;; prompt-reason : -> string
101
102 (define (prompt-reason)
103 (display "Reason (one line): ")
104 (flush-output)
105 (define line (read-line (current-input-port)))
106 (if (eof-object? line) "" line))
dynamic-wind takes three thunks: a before-thunk (here void), a body-thunk, and an after-thunk. The after-thunk runs whether the body completes normally or raises an exception, analogous to Python’s try/finally. This guarantees the three temporary files are cleaned up regardless of what goes wrong.
Colorizing the Diff
print-colored-diff walks each line of the unified diff output and applies ANSI terminal color codes:
1 (define (print-colored-diff diff-text)
2 (for ([line (in-list (string-split diff-text "\n"))])
3 (cond
4 [(color-enabled?)
5 (cond
6 [(or (string-prefix? line "+++")
7 (string-prefix? line "---")
8 (string-prefix? line "@@"))
9 (displayln (string-append ANSI-CYAN line ANSI-RESET))]
10 [(string-prefix? line "+")
11 (displayln (string-append ANSI-GREEN line ANSI-RESET))]
12 [(string-prefix? line "-")
13 (displayln (string-append ANSI-RED line ANSI-RESET))]
14 [else (displayln line)])]
15 [else (displayln line)])))
Lines beginning with + are added lines and appear green; lines beginning with - are removed and appear red; diff headers (+++, ---, @@) appear cyan. This makes it straightforward to review a proposed change without reading both full file versions. The color-enabled? parameter turns the codes off for --plain, --no-color, and piped output, where escape sequences would only corrupt a log.
The Approval Prompt
prompt-yes-no-skip reads a cooked line from stdin and re-prompts until it sees one of the three answers:
1 ;; prompt-yes-no-skip : -> (or 'yes 'no 'skip)
2
3 (define (prompt-yes-no-skip)
4 (let loop ()
5 (display "\nApply this change? [y]es / [n]o / [s]kip and tell the model why: ")
6 (flush-output)
7 (define line (read-line (current-input-port)))
8 (define norm (if (eof-object? line) "" (string-downcase (string-trim line))))
9 (cond
10 [(member norm '("y" "yes")) 'yes]
11 [(member norm '("n" "no")) 'no]
12 [(member norm '("s" "skip")) 'skip]
13 [else
14 (displayln "Please answer y, n, or s.")
15 (loop)])))
16
17 ;; ---------------------------------------------------------------------------
18 ;; prompt-reason : -> string
19
20 (define (prompt-reason)
21 (display "Reason (one line): ")
22 (flush-output)
23 (define line (read-line (current-input-port)))
24 (if (eof-object? line) "" line))
The prompt uses plain read-line rather than the readline-based reader that drives the REPL. Approval happens in the middle of a tool call, where a short y/n/s answer is easier to reason about than an edited line with history, and read-line behaves correctly when stdin is a pipe, which is what makes --stdin and one-shot scripting work. prompt-reason collects the explanation for a skipped edit.
Web Search Integration
search.rkt provides two search backends with identical return shapes, making them interchangeable at the call site. Here is the complete file:
1 #lang racket
2
3 (require net/http-easy
4 net/uri-codec
5 json
6 racket/string)
7
8 (provide brave-search
9 exa-search)
10
11 (define EXA-ENDPOINT "https://api.exa.ai/search")
12
13 ;; ---------------------------------------------------------------------------
14 ;; Brave Search
15 ;; Returns (listof (list url title description))
16
17 (define (brave-search query [num-results 5])
18 (define api-key (getenv "BRAVE_SEARCH_API_KEY"))
19 (unless (and api-key (not (string=? api-key "")))
20 (error 'brave-search "BRAVE_SEARCH_API_KEY environment variable not set"))
21 (define encoded (uri-encode query))
22 (define url (format "https://api.search.brave.com/res/v1/web/search?q=~a&count=~a"
23 encoded num-results))
24 (define headers
25 (hash 'X-Subscription-Token api-key
26 'content-type "application/json"
27 'accept "application/json"))
28 (define resp
29 (get url #:headers headers))
30 (define data (response-json resp))
31 (define web (hash-ref data 'web (hash)))
32 (define results (hash-ref web 'results '()))
33 (for/list ([r (in-list results)])
34 (list (hash-ref r 'url "")
35 (hash-ref r 'title "")
36 (hash-ref r 'description ""))))
37
38 ;; ---------------------------------------------------------------------------
39 ;; Exa AI Search
40 ;; Returns (listof (list url title highlight))
41
42 (define (exa-search query [num-results 5])
43 (define api-key (getenv "EXA_SEARCH_API_KEY"))
44 (unless (and api-key (not (string=? api-key "")))
45 (error 'exa-search "EXA_SEARCH_API_KEY environment variable not set"))
46 (define payload
47 (hash 'query query
48 'type "auto"
49 'numResults num-results
50 'contents (hash 'highlights #t)))
51 (define headers
52 (hash 'content-type "application/json"
53 'authorization (string-append "Bearer " api-key)))
54 (define resp
55 (post EXA-ENDPOINT
56 #:headers headers
57 #:json payload))
58 (define data (response-json resp))
59 (define results (hash-ref data 'results '()))
60 (for/list ([r (in-list results)])
61 (list (hash-ref r 'url "")
62 (hash-ref r 'title "")
63 (let ([hl (hash-ref r 'highlights '())])
64 (if (and (list? hl) (not (null? hl))) (first hl) "")))))
Both functions return a list of (url title description) triples. Brave uses a GET request with an API key header and returns web search results with title and description snippets. Exa uses a POST with a JSON body and returns neural search results with highlighted excerpts.
The net/http-easy package, installable via raco pkg install http-easy, provides the get, post, and response-json procedures used here.
The Main REPL
Provider Dispatch
agent.rkt is the entry point that ties everything together. It no longer holds a provider parameter of its own. The active provider is a profile in the harness config, and every model call goes through two dispatch functions that read it:
1 (define model-override (box #f))
2
3 (define (config-loaded?)
4 ;; Any harness config loaded at all?
5 (not (hash-empty? (harness-config))))
6
7 (define (active-provider-hash)
8 (and (config-loaded?) (config-active-provider)))
9
10 (define (require-provider who)
11 ;; -> provider hash, or a clear error when nothing is configured.
12 (or (active-provider-hash)
13 (error who "~a"
14 (string-append
15 "no active provider profile; define \"providers\" in "
16 "~/.coding_harness.json or .local_coding_harness.json"))))
17
18 (define (active-provider-type)
19 ;; -> 'mlx | 'openai (wire format of the active chat provider)
20 (define p (active-provider-hash))
21 (if p (provider-type p) 'openai))
22
23 (define (using-mlx?) (eq? (active-provider-type) 'mlx))
24
25 (define (current-provider-name-or-legacy)
26 (or (config-active-provider-name) "?"))
27
28 (define (current-model-id)
29 (or (unbox model-override)
30 (let ([p (active-provider-hash)])
31 (and p (provider-model p)))
32 "?"))
33
34 (define (set-current-model! m)
35 ;; /model <id> or --model: override the active profile's model this session.
36 (set-box! model-override m))
37
38 (define (switch-provider! name)
39 ;; Select a profile and drop any model override so the new profile's own
40 ;; model takes effect (provider selection is always applied before --model).
41 (define active (config-set-active-provider! name))
42 (set-box! model-override #f)
43 active)
44
45 ;; Plain (no-tools) chat and agentic (tool-calling) chat both dispatch on the
46 ;; active provider profile. Explicit keyword args win; otherwise generation
47 ;; parameters come from the profile, and a parameter the profile omits is left
48 ;; out of the request rather than defaulted in code.
49 (define (chat-provider-hash) (active-provider-hash))
50
51 (define (apply-mlx-provider! provider)
52 ;; Point the mlx module at the given profile (or clear).
53 (mlx-active-provider (or provider #f)))
54
55 (define (profile-gen provider key explicit)
56 (or explicit (generation-ref (provider-generation provider) key #f)))
57
58 ;; Plain (no-tools) chat
59 (define (llm-chat msgs
60 #:max-tokens [max-tokens #f]
61 #:temperature [temperature #f])
62 (define p (require-provider 'llm-chat))
63 (define mt (profile-gen p 'max_tokens max-tokens))
64 (define tp (profile-gen p 'temperature temperature))
65 (case (provider-type p)
66 [(mlx)
67 (parameterize ([mlx-active-provider p])
68 (mlx-chat msgs #:model-id (current-model-id)
69 #:max-tokens mt #:temperature tp))]
70 [else
71 (chat msgs #:model-id (current-model-id)
72 #:max-tokens mt #:temperature tp)]))
73
74 ;; Agentic (tool-calling) chat -- uses the same active provider as plain chat.
75 (define (llm-chat-with-tools msgs tools)
76 (define p (require-provider 'llm-chat-with-tools))
77 (case (provider-type p)
78 [(mlx)
79 (parameterize ([mlx-active-provider p])
80 (mlx-chat-with-tools msgs tools #:model-id (current-model-id)))]
81 [else
82 (chat-with-tools msgs tools #:model-id (current-model-id))]))
require-provider is the single place that decides whether the harness can run at all. If no config file declares any providers, it raises an error naming both config paths instead of silently falling back to a compiled-in default. active-provider-type reduces the profile’s type field to 'mlx or 'openai, and using-mlx? uses that answer for /tokens and the banner.
Both dispatch functions resolve the model id and the generation parameters the same way: an explicit keyword argument wins, otherwise the value comes from the profile’s generation block, otherwise the parameter is left out of the request entirely. For MLX the profile is bound into mlx-active-provider with parameterize around the call, which is how a module-level client learns which endpoint and model to use without reaching for a global variable.
A session-level model override sits in front of the profile’s model. /model <id> and --model <id> set it, and switch-provider! clears it so that changing providers always adopts the new profile’s own model. The /provider command with no argument prints the active profile and lists every profile in the config; with an argument it switches to that profile.
Intent Classification
Before sending any message to the model, agent.rkt classifies the user’s intent as one of three categories: "general", "coding", or "hybrid". The classification uses a two-stage approach.
Stage one is a keyword heuristic, free and instant:
1 (define GENERAL-KEYWORDS
2 (list "movie" "film" "cinema" "theater" "theatre" "showing" "playing" "showtime"
3 "weather" "forecast" "rain" "snow" "temperature outside"
4 "restaurant" "recipe" "menu" "where to eat"
5 "news" "sports" "score" "standings"
6 "near me" "nearby" "directions to"
7 "hotel" "flight" "travel" "vacation"
8 "population of" "history of" "capital of"
9 "who is " "who was " "where is " "when is " "when does "
10 "price of" "cost of" "how much does"))
11
12 (define CODING-KEYWORDS
13 (list ".lisp" ".py" ".js" ".ts" ".java" ".cpp" ".go" ".rb" ".rs" ".c "
14 "def " "class " "function " "refactor" "implement " "compile" "makefile"
15 "stacktrace" "segfault" "git commit" "git push" "git pull"
16 "unit test" "pull request" "fix the bug" "add a function" "write a function"))
1 (define (heuristic-classify lower)
2 (cond
3 [(for/or ([kw (in-list GENERAL-KEYWORDS)])
4 (string-contains? lower kw))
5 "general"]
6 [(for/or ([kw (in-list CODING-KEYWORDS)])
7 (string-contains? lower kw))
8 "coding"]
9 [else #f]))
for/or is the Racket comprehension form that returns the first “truthy” value or #f if none is found. The general list is checked first, so a query that mentions both a film title and a file extension is routed to the general path. The ordering of the two lists is the tie-breaker.
If the heuristic returns #f (the query is ambiguous), stage two calls the model with a minimal two-message conversation and requests a single-word answer:
1 (define (llm-classify user-line)
2 (with-handlers ([exn:fail? (lambda (e)
3 (displayln (format "[Classifier LLM error: ~a — defaulting to coding]" (exn-message e)))
4 "coding")])
5 (define msgs
6 (list (hash 'role "system" 'content "You are a one-word query classifier. Reply with exactly one word and nothing else.")
7 (hash 'role "user" 'content
8 (string-append
9 "Classify this query as exactly one word — GENERAL, CODING, or HYBRID:\n"
10 "GENERAL = factual or informational; nothing to do with writing, editing, or debugging code.\n"
11 "CODING = writing, editing, refactoring, or debugging code or files.\n"
12 "HYBRID = coding question that benefits from web docs or library references.\n"
13 (format "Query: ~a\n" user-line)
14 "One-word answer:"))))
15 (define raw (llm-chat msgs #:max-tokens 10 #:temperature 0.0))
16 (define up (string-upcase (string-trim raw)))
17 (cond
18 [(string-contains? up "GENERAL") "general"]
19 [(string-contains? up "HYBRID") "hybrid"]
20 [else "coding"])))
21
22 (define (classify-intent user-line)
23 (or (heuristic-classify (string-downcase user-line))
24 (llm-classify user-line)))
max-tokens 10 and temperature 0.0 keep the classifier call cheap and deterministic. Both are passed explicitly, and because the dispatcher prefers an explicit argument over the profile’s value, the classifier’s tiny budget is not overridden by a profile that declares a large max_tokens. If the classifier itself fails, the handler defaults to "coding", a conservative choice that enables the full tool set.
Routing to the Model
send-to-model uses the classification to choose the right system prompt and call path:
1 (define (send-to-model user-line)
2 (define intent (classify-intent user-line))
3 (define label
4 (hash-ref (hash "general" "web search, no coding tools"
5 "coding" "coding tools, no search"
6 "hybrid" "coding tools + web search if /search is on")
7 intent))
8 (unless (cli-quiet?)
9 (displayln (format "[intent: ~a → ~a]" intent label)))
10 (cond
11 [(string=? intent "general")
12 (define content (or (maybe-search user-line #t) user-line))
13 (define msgs
14 (list (hash 'role "system" 'content GENERAL-SYSTEM-PROMPT)
15 (hash 'role "user" 'content content)))
16 (define reply (llm-chat msgs))
17 (displayln (format "\n~a" (clean reply)))]
18 [(string=? intent "coding")
19 (define updated (append (unbox messages-box) (list (hash 'role "user" 'content user-line))))
20 (define-values (reply new-messages)
21 (llm-chat-with-tools updated ENABLED-TOOLS))
22 (set-box! messages-box new-messages)
23 (displayln (format "\n~a" (clean reply)))]
24 [else ; hybrid
25 (define content (or (maybe-search user-line #f) user-line))
26 (define updated (append (unbox messages-box) (list (hash 'role "user" 'content content))))
27 (define-values (reply new-messages)
28 (llm-chat-with-tools updated ENABLED-TOOLS))
29 (set-box! messages-box new-messages)
30 (displayln (format "\n~a" (clean reply)))]))
General questions use a lightweight one-shot call and a simple system prompt, and they always run a web search. Coding requests go through the full agentic tool loop using a system prompt that describes the five tools and the rules for using them. Hybrid requests get the tool loop plus web search results prepended to the message, but only when /search is on. The [intent: ...] line is suppressed in quiet mode so scripted runs stay clean.
The System Prompt
The coding system prompt is set once per session and injected as the first message with role "system". It tells the model which tools are available and how to use them correctly:
1 (define SYSTEM-PROMPT-TEMPLATE
2 "You are an interactive coding assistant working in the directory {cwd}.\n\nRules:\n- Use read_file, list_dir, and grep to understand the code BEFORE proposing edits.\n- To EDIT an existing file: read_file it first, then pass its exact current contents\n as `old` to propose_edit.\n- To CREATE a new file: call propose_edit with the empty string \"\" as `old` and\n the full desired contents as `new`. Do not call read_file first for a file that\n does not exist yet.\n- One file per propose_edit call. Keep diffs small and focused.\n- If the user rejects an edit or `make check` fails, ask for clarification instead\n of retrying blindly.\n- run_shell only accepts whitelisted commands: make, ls, pwd, cat, uv.\n- When you are done, reply with a short natural-language summary of what changed.")
3
4 (define GENERAL-SYSTEM-PROMPT
5 "You are a helpful assistant. Answer the user's question clearly and concisely using the web search results provided. Do not reference files, directories, or code editing tools unless the user explicitly asks about code.")
6
7 (define COMPACT-SYSTEM-PROMPT
8 "You are a context compactor for a coding assistant. Summarize the conversation transcript into a compact brief that will replace it. Preserve: the user's goals and instructions, decisions made, files created or modified (with paths), important code and tool-output details, and outstanding tasks. Write dense bullets, no preamble.")
The {cwd} placeholder is replaced with the actual working directory at session start. Telling the model the working directory helps it construct relative paths for read_file and list_dir calls. The prompt also states the rule that matters most in practice: read a file before editing it, and pass its exact current contents as old.
Context Management
As an agentic conversation grows, every tool result is appended to the message list, and the context window fills up. agent.rkt provides two commands to manage this. /context shows a formatted table of messages with estimated character and token counts, plus a short preview of each message:
1 (define (show-context)
2 (define msgs (unbox messages-box))
3 (define total (for/sum ([m (in-list msgs)]) (message-char-size m)))
4 (displayln "")
5 (displayln (format "Context: ~a message~a, ~a chars, ~a tokens (est.)"
6 (length msgs)
7 (if (= (length msgs) 1) "" "s")
8 total
9 (quotient total 4)))
10 (displayln "")
11 (displayln (format " ~a ~a ~a ~a"
12 (~a "#" #:width 3 #:align 'right)
13 (~a "role" #:width 9)
14 (~a "chars" #:width 7 #:align 'right)
15 "preview"))
16 (displayln (format " ~a ~a ~a ~a"
17 (make-string 3 #\-)
18 (make-string 9 #\-)
19 (make-string 7 #\-)
20 (make-string 50 #\-)))
21 (for ([m (in-list msgs)] [i (in-naturals 1)])
22 (define lines (wrap-preview (message-preview m)))
23 (displayln (format " ~a ~a ~a ~a"
24 (~a i #:width 3 #:align 'right)
25 (~a (hash-ref m 'role "?") #:width 9)
26 (~a (message-char-size m) #:width 7 #:align 'right)
27 (first lines)))
28 (for ([extra (in-list (rest lines))])
29 (displayln (format " ~a ~a ~a ~a"
30 (make-string 3 #\space)
31 (make-string 9 #\space)
32 (make-string 7 #\space)
33 extra))))
34 (displayln ""))
Each preview is collapsed to a single line and wrapped to at most three 60-character lines, with an ellipsis marking text that still does not fit, so one long tool result cannot flood the table. The token estimate divides the character count by four, which is a rough but useful approximation for English and for code.
/compact sends the whole transcript to the model with the compactor system prompt shown above, gets back a dense summary, and replaces everything except the original system prompt with that summary. The final context table is printed again so you can see the size drop. This trades a little fidelity for a lot of context budget, keeping the model inside its window on long sessions.
Skills
The agent supports loading “skills” from ~/.agents/skills/<name>/SKILL.md. Each skill file is a Markdown document that is injected into the conversation as a system message, telling the model to treat it as authoritative guidance. /skills lists the available skills, parsing a description: field from each file’s YAML frontmatter, and /<skill-name> loads one. This lets you package reusable instructions that the model will follow for the rest of the session.
Line Editing, History, and Completion
Interactive input is the job of line-input.rkt. When stdin is a terminal and Racket’s bundled readline collection can be loaded, the REPL gets cursor editing, in-session history with the arrow keys and Ctrl-R, and Tab completion, all in process and with no external rlwrap wrapper. Two things can go wrong, and both degrade to a plain read-line instead of failing the harness: the native library may be missing, and stdin may not be a terminal at all (a pipe, a heredoc, or --stdin).
The guarded load is the first piece:
1 (define (load-backend!)
2 ;; Instantiate readline/readline once, tolerating a missing module or a
3 ;; missing native library. All lookups happen before any assignment so a
4 ;; failure can never leave the module half-initialised.
5 (unless backend-tried?
6 (set! backend-tried? #t)
7 (with-handlers ([exn:fail? (lambda (_) (void))])
8 (define (get name) (dynamic-require 'readline/readline name))
9 (define rl (get 'readline))
10 (define add-history (get 'add-history))
11 (define history-length (get 'history-length))
12 (define history-get (get 'history-get))
13 (define set-completion (get 'set-completion-function!))
14 (define set-completion-append-char (get 'set-completion-append-character!))
15 (set! readline-proc rl)
16 (set! add-history-proc add-history)
17 (set! history-length-proc history-length)
18 (set! history-get-proc history-get)
19 (set! set-completion-proc set-completion)
20 (set! set-completion-append-char-proc set-completion-append-char))))
21
22 (define (terminal-stdin?)
23 ;; readline captures the input port when its module is instantiated, so only
24 ;; engage it when that port really is the terminal's stdin.
25 (and (terminal-port? (current-input-port))
26 (eq? 'stdin (object-name (current-input-port)))))
27
28 (define (line-input-available?)
29 ;; -> boolean. Loads the backend on first use, and only when stdin is a
30 ;; terminal, so piped/--stdin runs never touch libedit at all.
31 (and (terminal-stdin?)
32 (begin (load-backend!) (and readline-proc #t))))
dynamic-require is called inside with-handlers, and every lookup happens before any assignment, so a failure can never leave the module half-initialized. terminal-stdin? exists because readline captures the input port when its module is instantiated: the backend is engaged only when that port really is the terminal’s stdin, which is what keeps piped output byte-for-byte identical to the old behavior.
History is trickier than it looks. Editline does not add accepted lines to its own history while GNU Readline does, so the module probes once on the first non-empty line and fills the gap itself if the backend left the history unchanged:
1 (define history-probed? #f)
2 (define backend-adds-history? #f)
3
4 (define (remember-line! line before-count)
5 (define grew? (> (history-length-proc) before-count))
6 (cond
7 [(not history-probed?)
8 (set! history-probed? #t)
9 (set! backend-adds-history? grew?)
10 (when (and (not grew?) (not (string=? line "")))
11 (add-history-proc line))]
12 [(and (not backend-adds-history?) (not (string=? line "")))
13 (add-history-proc line)]))
Persistence goes to ~/.coding_agent_history. save-history! keeps the last max-history (1000) entries and writes them atomically, using negative history indices that count back from the newest entry, which sidesteps the zero-versus-one base difference between the two backends.
Completion is installed as a callback that receives the word under the cursor. Tab on a word beginning with / completes slash commands; elsewhere it completes provider profile names, search engines, and the model ids declared in the harness config:
1 (define (completion-candidates word)
2 ;; Readline completes the word under the cursor rather than the whole line, so
3 ;; the command set and the argument sets are merged and filtered by prefix.
4 ;; Tab in the middle of prose only reacts to words that actually begin a known
5 ;; name (command, provider profile, engine, or model id).
6 (define (matching xs) (filter (lambda (x) (string-prefix? x word)) xs))
7 (if (string-prefix? word "/")
8 (matching SLASH-COMMANDS)
9 (append (matching (config-provider-names))
10 (matching SEARCH-ENGINES)
11 (matching (configured-model-ids)))))
12
13 (define (setup-line-input!)
14 ;; Enable readline editing, persisted history, and Tab completion when stdin
15 ;; is a terminal and the readline backend is present. Returns #t when active.
16 (define available? (line-input-available?))
17 (when available?
18 (set-history-file! (default-history-file))
19 (install-completion! completion-candidates)
20 (plumber-add-flush! (current-plumber)
21 (lambda (_) (save-history!))))
22 available?)
setup-line-input! wires the three pieces together and registers a plumber flush hook, so history is saved when the process exits normally. It returns #t only when the backend is actually active, and the REPL works unchanged when it returns #f.
The REPL Loop
The main loop in agent.rkt is a straightforward tail-recursive function:
1 (define (run-repl)
2 (unless (config-loaded?) (load-harness-config))
3 (register-all)
4 (reset-conversation)
5 (setup-line-input!)
6 (print-banner)
7 (let loop ()
8 (define line
9 (with-handlers ([exn:fail? (lambda (_) eof)])
10 (read-input-line "\n> ")))
11 (cond
12 [(eof-object? line)
13 (save-history!)
14 (displayln "")
15 (void)]
16 [else
17 (define trimmed (string-trim line))
18 (cond
19 [(string=? trimmed "") (loop)]
20 [else
21 (define cmd (handle-slash-command trimmed))
22 (cond
23 [(eq? cmd 'quit) (void)]
24 [(eq? cmd 'continue) (loop)]
25 [else
26 (with-handlers ([exn:fail? (lambda (e)
27 (displayln (format "\nError talking to model: ~a" (exn-message e)))
28 (flush-output))])
29 (send-to-model trimmed))
30 (loop)])])])))
handle-slash-command recognizes /reset, /history, /context, /compact, /model, /provider, /debug, /search, /tokens, /help, /skills, and any other /<name> as a skill lookup, all before any model call is made. Input is read through read-input-line, so the same loop gets readline editing on a terminal and a plain prompt otherwise. The module+ main form lets agent.rkt be both loaded as a library (for testing) and run directly from the command line.
The Command-Line Interface
Because agent.rkt ends in a module+ main submodule, it is also an ordinary Unix-style command. If any prompt is supplied, the harness runs a single task and exits; otherwise it starts the REPL. The flags are declared with racket/cmdline:
1 (command-line
2 #:program "coding-agent"
3 #:once-each
4 [("--stdin") "Read prompt from stdin (pipe/heredoc)" (set! stdin? #t)]
5 [("-y" "--yes") "Auto-approve all propose_edit diffs (still shows diff)" (set! yes? #t)]
6 [("--dry-run") "Show diffs but do not write files" (set! dry-run?flag #t)]
7 [("--provider") prov "LLM provider profile (from harness config) or fireworks/mlx/omlx/sushi" (set! provider-str prov)]
8 [("--model") mid "Model id for current provider" (set! model-str mid)]
9 [("--cwd") dir "Working directory before running" (set! cwd-str dir)]
10 [("--debug") "Enable debug logging (same as /debug)" (set! debug?flag #t)]
11 [("-q" "--quiet") "Quiet: no banner, no [intent] line, less tool chatter" (set! quiet?flag #t)]
12 [("--plain" "--no-color") "Plain output: no ANSI colors in diffs" (set! plain?flag #t)]
13 [("-v" "--version") "Show version and exit" (set! version?flag #t)]
14 #:multi
15 [("-p" "--prompt") p "Prompt text (repeatable, joined with newlines)" (set! prompt-parts (append prompt-parts (list p)))]
16 #:args args
17 (set! positional-parts args))
--help is generated by racket/cmdline and exits with status 0. --version prints the version and exits before any provider check. A prompt can arrive as positional arguments, as one or more -p/--prompt values, or on stdin with --stdin; build-prompt joins every source with blank lines and trims the result, so mixing them works.
Two flags exist for automation. --yes prints the diff and applies it without prompting, and --dry-run prints the diff and writes nothing. --quiet suppresses the banner, the [intent] line, and most tool chatter, while --plain (or --no-color) turns off ANSI codes, which is what you want when redirecting output to a file.
Exit Codes
One-shot runs are meant to compose with shell scripts, so the harness maps outcomes to exit codes: 0 for a normal finish, 1 for a model or network error, 2 when a make check gate failed, 3 when the user rejected or skipped a change, and 5 for bad arguments. The infer-exit-code helper derives the code by scanning the tool results in the conversation for the strings the tools produce, which keeps the decision in one place instead of threading a status value through the agentic loop.
Running the Agent
Installation
Install Racket 8.11 or later from racket-lang.org. Then install the http-easy HTTP client package:
1 raco pkg install --auto http-easy
Export your API keys. Which key you need depends on the provider profiles you configure; the Fireworks profile used in this chapter reads FIREWORKS_API_KEY, and the search keys are optional:
1 export FIREWORKS_API_KEY=fw_...
2 export BRAVE_SEARCH_API_KEY=...
3 export EXA_SEARCH_API_KEY=...
Finally, create ~/.coding_harness.json with at least one provider. A local MLX profile needs no key, only a running server:
1 mlx_lm.server --model mlx-community/gemma-4-26B-A4B-it-OptiQ-4bit --port 11434
With no providers configured the harness refuses to start and prints both config paths, which is the intended failure mode: there is no compiled-in provider to fall back to.
Starting the REPL
1 make run
or directly:
1 racket agent.rkt
The banner shows the working directory, the active provider profile, and the active model:
1 Coding Agent REPL. /help for commands, /quit to exit.
2 cwd: /Users/mark/myproject
3 provider: fireworks
4 model: accounts/fireworks/models/deepseek-v4p1-flash
Sample Session
The following session asks the agent to add a helper function to an existing file. Lines beginning with > are user input; everything else is agent output.
1 > add a function called word-count that takes a string and returns the number of words
2
3 [intent: coding → coding tools, no search]
4 * read_file utils.rkt
5 * propose_edit utils.rkt
6
7 --- a/utils.rkt
8 +++ b/utils.rkt
9 @@ -14,3 +14,7 @@
10 (define (trim-lines text)
11 (string-join (map string-trim (string-split text "\n")) "\n"))
12 +
13 +(define (word-count str)
14 + (length (string-split str)))
15 +
16 +(provide word-count)
17
18 Apply this change? [y]es / [n]o / [s]kip and tell the model why: y
19 applied; make check passed
20
21 Added `word-count` to utils.rkt. It splits the string on whitespace using
22 `string-split` (which treats consecutive spaces as one separator) and returns
23 the length of the resulting list.
24
25 > /tokens
26
27 Session token usage:
28 Prompt tokens: 1842
29 Completion tokens: 87
30 Total tokens: 1929
31 Estimated cost: $0.000282 ($0.1400/M input, $0.0280/M cached input, $0.2800/M output)
32
33 > /quit
Enabling Web Search
Toggle search on with /search. Switch between engines with /search brave or /search exa:
1 > /search brave
2 Web search ON (engine: brave)
3
4 > what is the current version of Racket?
5
6 [intent: general → web search, no coding tools]
7 [Web search results for: what is the current version of Racket?]
8 1. Racket -- A programmable programming language
9 https://racket-lang.org
10 Racket 8.14 was released on ...
11 ...
12
13 As of mid-2026, the current stable release of Racket is version 8.14.
Switching Providers
The /provider command with no argument reports the active profile and everything the config declares. With a profile name it switches, and the next call uses that profile’s endpoint, model, and generation settings:
1 > /provider
2 Current provider: fireworks (model: accounts/fireworks/models/deepseek-v4p1-flash)
3 Available profiles: deepseek, fireworks, mlx, omlx, sushi
4
5 > /provider mlx
6 Provider set to profile 'mlx' (model: mlx-community/gemma-4-26B-A4B-it-OptiQ-4bit)
7
8 > /tokens
9
10 Session token usage (local MLX -- no API cost):
11 Prompt tokens: 0
12 Completion tokens: 0
13 Estimated cost: $0 (local model mlx-community/gemma-4-26B-A4B-it-OptiQ-4bit)
Running One-Shot Commands
Any prompt makes the run non-interactive. This is the form to use from a script or a Makefile:
1 racket agent.rkt -p "add a docstring to word-count"
2 racket agent.rkt --stdin --quiet --plain < task.txt > out.txt
3 racket agent.rkt --dry-run -p "rename foo to bar"
4 racket agent.rkt -y -p "fix the failing test"
5 echo $? # 0 ok, 2 make check failed, 3 rejected, 5 bad args
Building an Executable
The Makefile also builds and installs a standalone binary, and it regenerates shell completions:
1 make make-executable # builds ./coding-agent
2 make install PREFIX=/usr/local # copy to $PREFIX/bin
3 make completions # bash/zsh/fish into ./completions/
4 make dist # raco distribute -> ./dist/
The executable runs from any directory on this machine. Because it is built with ++lib readline/readline, line editing survives the packaging step, while a machine without libedit still falls back to plain input at run time.
Interpreting the Output
When the agent prints * tool-name arguments it is showing a tool call in progress. The tool name and a truncated version of the arguments help you follow the model’s reasoning. read_file utils.rkt means the model decided it needs to see the file before editing it, a sign it is following the system prompt rules. propose_edit always appears after a read_file for the same path.
make check passed tells you both that the model’s proposed syntax was valid Racket and that your project’s own compile step accepted it. If you see make check FAILED, the failure output follows immediately and appears in the agent’s next prompt, giving the model a second chance to correct the error autonomously.
The /tokens output shows prompt tokens growing much faster than completion tokens. That is expected in an agentic loop: the conversation history (including tool results, which can be long) is re-sent to the model on every iteration, while the model’s replies are comparatively short. The cost estimate uses the rates declared in the active provider’s pricing block, and a profile without one prints n/a (no "pricing" block for this provider) instead of guessing. A local MLX profile always reports $0.
The [intent: ...] line tells you how the agent routed your request. A general question is answered without touching any tools; a coding task gets the full tool loop. If the routing looks wrong, you can inspect the keyword lists and adjust them. The line is hidden under --quiet.
A line such as (stopped: the model repeated the identical tool call(s) 2 times ...) means the repetition guard fired. That is a signal that the model is too weak for the task or is emitting malformed arguments, not that the harness has crashed.
Wrap Up
This chapter built a complete Racket coding agent in roughly 2,800 lines across nine focused modules. The main ideas were:
Separation of concerns. The shared agentic loop, the two provider clients, the configuration layer, the tool registry, the approval UI, the search backends, and the line editor each live in their own file. The only module that cannot be required statically is the line editor, and it isolates that fact: it loads readline with dynamic-require and falls back to plain input, so a missing native library is a degradation rather than a failure.
Provider abstraction. The agentic loop in chat-loop.rkt is parameterized by a post-fn adapter, so Fireworks (cloud, SSE streaming) and MLX (local, OpenAI-compatible) both run through the identical loop. The only differences are in the transport at the boundary.
Configuration over compilation. Endpoints, models, generation parameters, API key variable names, and prices live in a JSON config with two merge layers. Adding a provider or changing a rate is a config edit, and no provider name appears anywhere in the Racket source.
The stale-base guard in propose_edit. Requiring the model to supply the exact current contents of a file before any edit is accepted prevents silent overwrites when the file changes between the read and edit steps. The mismatch is reported as a tool result the model can read and respond to.
Two-stage intent classification. A free keyword heuristic handles the common cases and falls back to a cheap model call only for ambiguous queries. Defaulting to "coding" on classifier failure keeps the full tool set available.
The make check feedback loop. Every accepted edit is immediately verified by the project’s own build target. Failures go back into the conversation history, giving the model the information it needs to self-correct on the next iteration.
A repetition guard in the loop. Small local models sometimes re-issue the identical failing call. Tracking the last few call signatures and stopping when one repeats turns a wasted run into an immediate, legible explanation.
These patterns (tool registries, approval gates with stale-base guards, intent routing, quality gates, provider adapters, and config-driven backends) apply broadly across languages and LLM providers. The Racket implementation here serves as a concrete reference for how each piece fits together at the system level.
Optional Practice Problems
Problem 1: Add a write_file tool
The agent currently has no way for the model to create a file without going through propose_edit. Add a write_file tool to tools.rkt that accepts a path and content parameter, writes the content directly (without a diff prompt), and returns a confirmation string. Register it in register-all and add it to ENABLED-TOOLS. Consider what safety constraints, if any, should prevent the model from overwriting files outside the working directory, and whether the existing hidden-file? predicate should apply.
Problem 2: Extend the shell whitelist dynamically
SHELL-WHITELIST is currently a compile-time constant. Add a /allow-cmd slash command to agent.rkt that lets the user append a command to the whitelist at runtime, so that /allow-cmd git would let the model run git status and git diff. Update handle-slash-command to recognize the new command and update the set stored in tools.rkt. Think about where the mutable whitelist state should live and how tools.rkt should expose it.
Problem 3: Persistent session history
At present, /reset discards the conversation history and there is no way to resume a previous session. Add two slash commands: /save <filename> that writes the current messages-box contents to a JSON file using jsexpr->string, and /load <filename> that reads that file and restores the conversation. Use string->jsexpr for loading. Handle file-not-found and malformed JSON gracefully by printing an error and leaving the existing history unchanged.
Problem 4: Token-budget guard
chat-with-tools* will keep iterating until the model stops calling tools or max-iterations is reached. Add a token-budget guard that checks the running token total after each iteration and returns early with a warning message if it exceeds a configurable threshold. The counters live in private boxes in fireworks-ai.rkt, and mlx-serve.rkt keeps a smaller pair of its own, so you will need to export a reader or thread a check through the post-fn adapter. Expose the threshold as a /budget <n> slash command that sets it, and a /budget command with no argument that prints the current setting and the remaining budget.
Problem 5: Second search backend – DuckDuckGo
Add a ddg-search function to search.rkt using the DuckDuckGo Instant Answer API at https://api.duckduckgo.com/?q=QUERY&format=json. The response contains a RelatedTopics array of objects with Text and FirstURL fields. Return results in the same (url title description) triple format as brave-search and exa-search. Update agent.rkt to accept /search ddg as a valid engine selection, and add it to the SEARCH-ENGINES list so Tab completes it.
Problem 6: Colored intent label in the REPL prompt
The line [intent: coding → coding tools, no search] is printed in plain text. Use ANSI codes to color the label: green for "coding", cyan for "general", and yellow ("\033[33m") for "hybrid". Update send-to-model in agent.rkt to apply the color. approval.rkt already defines the constants, but they are not in its provide list, so decide whether to export them, redefine them locally, or move them to a shared ansi.rkt module.
Problem 7: Retry on make check failure
Currently, when propose_edit runs make check and it fails, the failure output is returned to the model as a tool result, but the model must then propose a new edit from scratch. Modify tool-propose-edit so that on a make check failure it offers the user a [r]etry option at the approval prompt (in addition to the existing y/n/s choices). On retry, revert the file to its previous contents using call-with-output-file, print a confirmation, and return a result string that tells the model the file was reverted and includes the check output so it can try again with a corrected edit.
Problem 8: Stream the SSE deltas to the terminal
The Fireworks client already reassembles the SSE stream into a single response via parse-sse-response, but the user does not see the reply being written in real time. Modify post-fireworks so that, while parse-sse-response is accumulating the response, each delta.content fragment is also displayed to the terminal as it arrives. Consider how to do this without double-printing the final text (which the REPL also prints after the call returns), and whether the tool-call argument fragments should be hidden.
Problem 9: Add a provider profile without touching code
The provider layer is data, not code. Confirm that by adding a second local profile to ~/.coding_harness.json that points at a different port, for example oMLX on 8000 or sushi on 12345. Switch to it with /provider <name>, check it with /provider alone and with /tokens, and verify that Tab completion offers the new profile name. Then override one field in .local_coding_harness.json and confirm that the deep merge replaces just that field while the rest of the global profile survives.