← Markwork · Learning notes · October 4–5, 2026

Learning to build an agent from Markwork

In two days, the Markwork assistant grew from a chat panel with 4 read tools into an agent with 16 tools. It can now plan work, propose batches of changes, check the result, and pick the work back up after an interruption. One principle has been held since the first commit: the agent may read freely, but every file change must be approved by the user first. The last chapter shows the next step: Markwork assigns kanban cards to another agent (pi) that works in a separate project folder, while Markwork only holds the board. This page explains how that principle is kept at every stage. It covers data structures, the loop flow, safety limits, costs, and small simulations you can try yourself.

Tools offered to the model
16
from 4 · f6a4cc7 → 8d0cdec
Assistant tests (unit + GUI)
212
from 65 · all without a paid API
Code in src/agent/
3,726 lines
from 1,099 lines in 8 files → 21 files
Work folder writes without approval
0
all through Change + preflight; the external harness only in another project folder (Chapter 9)
The big picture

The journey: from answering to finishing the work

Each bar is one commit that changed the agent. The colors tell three kinds of tools apart: read (browsing files and Git history), proposal (producing changes that wait for approval), and control (the plan and verification, which do not touch files). Hover over a bar to see the number of tests and lines of code at that point.

ReadChange proposalsWork control
  1. f6a4cc7Chapters 1, 2 · The DeepSeek Assistant panel: context within a token budget, a stable prefix for the cache, BM25, an agent loop with 4 read-only tools.
  2. 2d0150fChapter 3 · Conversation history is saved as Markdown in .markwork/chats and can be resumed.
  3. 49021f3Chapter 3 · The prompt and identity are aligned with the personal workbench vision, no longer specific to book manuscripts.
  4. 3f387e6Chapter 3 · search_manuscript → search_documents; the tool descriptions are neutralized.
  5. 5a64789Chapter 4 · The first write tools (create_file, edit_file), which only produce proposals and wait for Apply/Reject.
  6. c6f71daChapter 4 · Proposals are reviewed in the same diff window as Git history.
  7. 8db7c55Chapter 4 · Fix for the panel being cut off after a long answer (scrolling through an idle).
  8. deb8aa5Chapter 4 · edit_kanban and multi-hunk diffs.
  9. 6c47266Chapters 5, 6 · The plan (set_work), batches (propose_batch), verification, the action journal, retry, and the history summary.
  10. 1c90735Chapter 7 · The Agent log window for monitoring rounds, reasoning, tools, and their results.
  11. 90c28bbChapter 8 · Insert, delete, move, replace all, partial approval, Undo, Git history tools, and structure verification.
  12. b1ab712Chapter 9 · The harness orchestrator: @pi cards are worked on by pi in another project folder, a queue per folder, the pi log, cards move to In Progress/Review.
  13. 0db63d2Chapter 9d · pi through RPC mode: a ⏸ waiting status when pi asks or an extension requests permission, answered from Markwork; steering while working and replies that continue the session.
  14. 8d0cdecChapter 9f · A [[note]] on a card becomes pi prompt context: the note contents from the work folder are copied into the prompt, limited per note and in total.
  15. b31556dChapter 5d · The set_work plan is shown as a checklist (WorkList): progress, a spinner on the step being worked on, turn duration; tidied to GNOME idioms in 4a35bc8 (Adw.Spinner, accessible labels).
  16. a8dd69aChapter 5d · The Assistant message box: the send button inside the message frame (6037ab4), the context summary appears immediately when the panel is open from the start.
  17. 9d6ea05Chapter 8d · The Git tools always answer: a watchdog in runGit ends a run whose result the runtime lost, or a git that hangs, as a failure the model can read.
  18. 2dbf902Four rules · The conversation life cycle (send, stop, folder-change guards, saving and reopening history, undo) moves from the panel into agent/chatcontroller.ts; the panel only draws through ChatView. The workspace files now come from WorkspaceRepository (e4e3d91), with an explicit freshness per read.
  19. 84dccddFour rules · The Assistant panel is split into its parts: settings, history, the context popover, and the proposal cards with their review window and Undo (ui/proposalcard.ts). Earlier commits split ChatSession.ask into a TurnRun with one method per tool (76acf0d), and planChange/planKanban into handlers per tool and action (1caf08a); every function is now at most complexity 15.
  20. 1bac3b8Chapter 10 · The assistant writes the commit message: one request without tools from the checked files' diffs within a 30,000-character budget, streamed into the commit box, which the user still reviews and commits.
Before adding capabilities

Four rules kept in every commit

The agent's capabilities quadrupled, but its basic shape stayed the same. Every new feature has to satisfy these four rules. That is why each chapter below can also be read as a way of satisfying them.

Pure

The agent logic in src/agent/ does not import GTK. Tools work on a SourceFile[], return values, and do not throw. That is why everything can be tested without a window and without a network.

Proposals, not writes

Tools that change things only produce a Change { before, after }. The window is what writes to disk, and that happens after the user presses Apply and the file contents are checked again.

Transparent

The user can see the context that was sent together with its token count, the browsing steps, the diff, the journal, and now the Agent log too. No decision is hidden.

Bounded

Every resource has a limit: context tokens, model rounds, the read budget, the size of tool results, the number of actions per batch, and paths inside the work folder.

agent/* (pure)→ used by agent/chatcontroller.ts→ ChatView ui/chat.ts→ asks for a decision the user→ Apply window.ts: preflight + write

The direction of calls is never reversed. agent/ knows nothing about the editor or widgets. The window only gives it a way to read files (ChatHost) and a way to ask for the user's decision (onProposal, onBatchProposal). Since 2dbf902 the turn itself (sending, stopping, refusing proposals after a folder change, saving, undo) is run by ChatController, and the panel only draws what ChatView tells it, so these paths are tested without a window.

Chapter 1 · commit f6a4cc7 · src/agent/context.ts

The model only knows what is sent: building context within a budget

The problem. The model cannot see the editor. If the work folder holds hundreds of KB of notes, it is impossible to send everything, and sending too much is also expensive and slow. So the quality of the answers is largely determined by a single function: buildContext().

The design. The context is split into two parts with different properties, and then ordered so that the part that rarely changes comes first:

system (stable) · instructions + the project map + the active document with line numbers. This part stays the same as long as the manuscript does not change, so its prefix can be served from the DeepSeek cache.
history · earlier questions and answers, at most 25% of the budget (HISTORY_SHARE).
note (changes) · the selection, the cursor position, files that were @mentioned, and relevant excerpts from other files. This part is attached only in front of the latest question.

Budget and priorities

DEFAULT_BUDGET = 48_000            // tokens
estimateTokens = s => ceil(s.length / 3)
// deliberately wasteful for Indonesian text
// order: selection → active document → @mention
//        → project map → BM25 excerpts

Excerpts and search

splitChunks(file, text)   // per heading,
  // sections >1,800 characters are split at blank lines
rankChunks(chunks, query) // BM25, k1 = 1.5, b = 0.75
  // headings count twice
buildQuery: question ×1, selection ×0.5,
            last 2 questions ×0.4

Earlier questions also go into the query with a small weight. That way, a follow-up question such as "then who handles it?" still finds the right excerpts. Suffixes are stripped roughly (the Indonesian -nya, -kan, -an, …) so that "rapatnya" matches "rapat".

Simulation 1: which excerpts go into the note? Interactive · real computation

This is a mock work folder with meeting notes, plans, research, and a task board. The heading splitter, tokenizer, and BM25 below are copies of the context.ts logic. Change the question or the budget to see which excerpts get selected. Excerpts that do not fit in the budget are not sent. The model can still read them later with a tool.

The budget is counted with the same estimate (3 characters ≈ 1 token). In Markwork, excerpts compete with the selection, the active document, and the project map within a 48,000-token budget. Here the budget is deliberately small so that the effect is visible.

Why the order matters: the prefix cache

Providers such as DeepSeek keep a cache for a message prefix that is exactly the same as in the previous request. A single different character at the start is enough to break the cache for the whole message after it. That is why the cursor position and the selection (which change with every question) must not be placed in system.

Simulation 2: tokens from the cache over one session Interactive · model

Compare two arrangements. Left: the cursor position is slipped into system. Right: Markwork's arrangement, with every part that changes going into the note at the end. A bar shows the input tokens per question. Blue marks tokens that can be served from the cache.

From cacheReprocessed
Cursor in system
Cursor in note (Markwork)
Cursor in system: from cache
–
Cursor in note: from cache
–

This model assumes the manuscript does not change, the note is ±2,500 tokens, and each turn adds ±600 tokens of history. In the left arrangement, the cursor position is written right after the instructions (±1,500 tokens), so only the instructions stay the same as in the previous request. This is not a DeepSeek measurement. The real numbers per round (input tokens, from cache, output) can be seen in the Agent log. When the manuscript changes, the active document in system changes too and the cache really does have to be rebuilt.

LessonContext is the main interface to the model. Split the context by how often its contents change, give each part a budget, and show the breakdown to the user. In Markwork, the Context button shows each part along with its token count.
Chapter 2 · commit f6a4cc7 · src/agent/session.ts, tools.ts

A bounded agent loop: the model asks, the code runs

A turn starts by calling the model. If the model asks for a tool, that tool is run, the result is sent back, and then the model is called again. The loop stops when the model gives an answer. The first version had only 4 read tools: list_files, search_manuscript (BM25), search_text (exact text), and read_file.

The core of the loop (condensed from f6a4cc7)

for (round = 0; round < MAX_ROUNDS; round++) {   // 10
  useTools = round < MAX_ROUNDS - 1
  result = await provider.chat({ messages,
             tools: useTools ? TOOLS : undefined })
  if (!result.toolCalls.length) break  // the answer
  for (call of result.toolCalls) {
    out = runTool(call.name, call.arguments, files)
    if (cost(out) > toolBudget) out = 'Budget exhausted…'
    messages.push({ role: 'tool', content: out })
  }
}

Three limits that guarantee the loop ends

  • Rounds: at most 10. The last round is called without tools, so the model is forced to answer.
  • Read budget: the total of tool results per turn ≤ 1× the context budget (TOOL_SHARE). If it is exceeded, the model is asked to answer with the information it already has and to say what has not been checked.
  • Result per tool: ≤ 6,000 tokens (MAX_RESULT_TOKENS). A longer result is cut off with a hint to read a specific range.

There are several small decisions that matter. Tools never throw: broken JSON arguments become an error message the model can read, so the model can fix it by itself. The result of read_file is given the same line numbers as the context, so "file:line" references are consistent. Tools are only offered if the user allows other files to be read (the project context switch). In addition, unsaved editor contents replace the version on disk.

Simulation 3: one agent turn, step by step Interactive · walkthrough

The user's request: "The launch meeting moved from Wednesday to Thursday, October 9; update all the documents." This is a walkthrough of the current loop flow (Chapters 5 and 8) with a scenario that has been written in advance. A real model may choose other steps. When the batch is proposed, you are the one who decides.

Model rounds
0 / 10
tools offered: –
Read budget remaining
48,000
Work status
–
files written: 0
LessonA reliable agent is not one that is "smart enough to stop by itself", but one whose loop has hard limits in code: the number of rounds, the read budget, and the size of results. Error messages and "budget exhausted" messages are written to be read by the model, not by the developer.
Chapter 3 · commits 2d0150f, 49021f3, 3f387e6 · chatstore.ts, transcript.ts, context.ts, tools.ts

Memory a human can read, and names that do not mislead

3a. A conversation as a Markdown file

Each conversation is saved as one file in <folder>/.markwork/chats/, with frontmatter (title, model, created) and then ## You / ## Assistant sections. The format is deliberately readable and editable by humans, not a hidden database.

What is saved

Questions and answers. Starting in Chapter 5, also work (the goal, steps, verification) and actions (the journal along with before/after snapshots). The .markwork/ folder automatically gets a .gitignore containing *, because its contents include document excerpts.

What is deliberately not saved

The model's thinking process, the context breakdown, the full read results, and token usage. The document context is rebuilt every turn (file contents may change), so the history stays small.

A trap that is avoided: the agent must not browse its own history as a document. WorkspaceRepository does not read dot folders, so .markwork/ never shows up in search_documents results. Without this rule, an old answer could be read as a "fact" from a document.

3b. Tool names are part of the prompt

Markwork was originally designed for writing books, so its prompt and tools talked about the "manuscript". After the vision was widened into a personal workbench (notes, meetings, research, plans), the name search_manuscript was changed to search_documents and all the tool descriptions were neutralized. The model picks tools by their names and descriptions. A name that is too narrow makes the model hesitate to use the tool for meeting minutes.

LessonTreat tool names and descriptions like interface text for a very literal user. When the scope of the product changes, re-check the prompt and the tool descriptions, not just the UI.
Chapter 4 · commits 5a64789, c6f71da, deb8aa5, 8db7c55 · src/agent/changes.ts, ui/proposalviewer.ts

From answering to proposing: changes as data

The need. Users want the agent not just to say "replace line 12", but to do it, without losing control over their files. The design. A write tool writes nothing. It only validates the arguments, and then computes the result as a value:

Data structure

interface Change {
  kind: 'create' | 'edit' | 'delete' | 'move'
  file: string     // relative path in the work folder
  before: string   // contents at the time of the proposal
  after: string    // contents if applied
  reason: string   // the reason, shown in the review window
  to?: string      // move only
}

The kind values delete/move were only added in Chapter 8. The first version only had create and edit.

Separating the tools

TOOLS          // read-only, run directly
CHANGE_TOOLS   // only produce a Change
// session:
change = planChange(name, args, files)
answer = await handlers.onProposal(change)
// ↑ a Promise that waits for the review window
if (answer.applied) applyToFiles(files, change)

Without an onProposal handler, the tools that change things are not offered at all.

edit_file asks for an old_text that appears exactly once. If it does not match, the model gets a message telling it to read the file first or to add surrounding lines. This approach is stronger than line numbers, because models often miscount lines. edit_kanban (deb8aa5) does not edit text directly. It uses the same board model as the kanban UI (parseBoard → operation → serializeBoard), so the result is always a valid board.

Preflight: the file contents must still be what they were at the proposal

There is a gap between the moment a proposal is made and the moment the user presses Apply. During that gap the user may well edit the file. That is why, before writing, the window calls preflight(). A proposal is rejected if before is no longer the same as the actual contents, instead of overwriting the user's edit.

function preflight(c, read) {
  const at = read(c.file)                  // null = does not exist
  if (c.kind === 'create') return at === null ? null : `${c.file} already exists`
  if (at === null)        return `${c.file} no longer exists`
  if (at !== c.before)    return `${c.file} changed since it was proposed; ask for a new proposal`
  if (c.kind === 'move' && read(c.to) !== null) return `${c.to} already exists`
  return null
}

Simulation 4: Apply, edit, Undo Interactive · real logic

The agent proposes the change below. Try Apply, and then Undo. After that, repeat but edit the file contents in the text box before pressing the button. preflight, changeState, and invertChange here are copies of those in changes.ts.

Actual contents of plans/october.md (you may edit it)
The edit_file proposal

changeState: –
The agent proposal review window with a red-green diff and Reject/Apply buttons
The review window (c6f71da) uses the same diff view as Git history, so users only need to learn one visual language.
A UI bug along the way (8db7c55). After a long answer, the Assistant panel was shown cut off and its footer disappeared. The cause: set_value() was called inside the adjustment's changed signal. The value changed, but the GTK viewport did not apply it. Scrolling is now scheduled through an idle. The test checks the position of the footer in the visible area, not only the adjustment value. A test that only checks internal numbers can pass even though the display is wrong.
LessonSeparate "deciding a change" from "writing a change". If a proposal is pure data with before and after, you get the diff, validation, conflict checking, and later Undo, all from the same structure.
Chapter 5 · commit 6c47266 · work.ts, batch.ts, verification.ts

Plan, batch, and verification: "done" must be checkable

The symptom. For multi-step work, such as moving a schedule across three documents and a task board, the model proposed changes one by one, sometimes forgot a file, and then declared "all done". The user had to approve five separate windows without seeing the whole picture.

5a. A plan that is stored: set_work

interface WorkState {
  goal: string
  status: 'running' | 'paused' | 'failed' | 'complete'
  steps: { text: string; status: 'pending' | 'done' | 'blocked' }[]   // 1–20
  note: string
  actionStart?: number          // journal index when this work started
  verification?: Verification   // the latest check result
}
running→ turn ends →paused · changes applied →running(verification cleared) · verification passes + all steps done →complete · error →failed

One important rule: the model cannot set complete itself. That status can only be reached if verify_work passes and all the steps have the status done. Every change applied afterwards clears the verification result, so the check has to be repeated. If, in the next turn, the file contents turn out no longer to match the changes that were applied, the complete status drops back to paused.

5b. Batches: many actions, one review

planBatch(args, files) {
  virtual = copy(files)          // a trial world
  for (item of actions) {        // ≤ 20
    c = planChange(item.tool, item.arguments, virtual)
    // delete/move lock the file
    merge per file: first before, last after
    applyToFiles(virtual, c)     // the next action
  }                              // sees this result
}
applyBatch(changes, host) {
  // 1. preflight ALL first; one failure = cancel everything
  // 2. write in order
  // 3. if a write fails: roll back what was written,
  //    return the error + any rollback failures
}
// an in-process transaction, no atomic guarantee
// against power loss

The actions in a batch are planned on a virtual copy, so a second action on the same file sees the result of the first. The final results are merged into one Change per file. This decision later makes the partial approval in Chapter 8 possible.

5c. Deterministic verification

verify_work does not ask another model for an opinion. The model states its criteria, and then code checks them against the actual contents: present / absent (exact text), kanban (a card is in a given list with a given status), and since Chapter 8 also structure. The session also adds checks of its own: every file that was changed but is not mentioned in the criteria automatically produces "Changed files not yet checked: FAILED". That way, the model cannot get through just by checking the easy files.

The Assistant panel showing a finished, verified work plan after the task board was updated
The plan, the approved batch, and the verification are shown in the panel. The status "done and verified" only appears after the checks pass.
LessonDo not trust the model's claim of "done". Make done a status that can only be reached through a deterministic check of the actual result, and revoke that status as soon as the result changes.

5d. A visible plan: one WorkState, two readers · commits b31556d, 4a35bc8 · ui/worklist.ts

The symptom. The plan had been stored since 6c47266, but the panel showed it as the workText() text as it was: - [x] Read the meeting decisions, exactly as sent to the model. While the agent worked, the user could not see at a glance which step was in progress, and the same plan appeared twice because the tool step line "Work plan → …" printed it as well.

// before: one text for the model and for humans
answer.work.set_text(workText(work))
// - [x] Read the meeting decisions
// - [ ] Check the plan and the task cards
// after: the model still receives workText(),
// the panel draws from the same data
answer.work.update(work, running, seconds)
current = running && work.status === 'running'
  ? steps.findIndex(s => s.status === 'pending') : -1

There is no new status. WorkList only reads the existing WorkState, so a turn that is running and a conversation reopened from .markwork/chats look the same. The only piece of information that does not come from the model is the step in progress: the first pending step, only while the turn is running. As soon as the turn finishes, the spinner disappears and the status goes back to the saved status (paused, failed, or done and verified), so the display never claims progress that was not recorded. The "Work plan" line in the tool step list is hidden when the plan is valid; a plan rejected by parseWork still appears as an error line.

The display follows GNOME idioms (see AGENT.md, "The UI must follow GNOME standards"): symbolic icons, the @success_bg_color/@success_fg_color color pair, Adw.Spinner when libadwaita ≥ 1.6 with a Gtk.Spinner fallback, and every row has the list item role with an accessible label "text — status", so that the status is not conveyed through color alone.

The Assistant panel showing the plan as a checklist: three steps with green checks, 3/3 steps, done and verified
The plan from the work in 5b–5c after the turn finished: progress, the saved status, and the turn duration under the goal; one check per step.

Simulation 9: from WorkState to a checklist Interactive · real logic

Change the status of each step and whether the turn is still running. The marker rules and the meta row are copies of the logic in WorkList.update() and stepRow(). Below it is the workText() text that the model receives for the same state.

The spinner appears only when the turn is running and the saved status is running. In Markwork the status complete can only be reached through verification (5c); here it can be chosen directly to see how it looks.

LessonThe same state may have two views: a compact text for the model, a widget for humans. Draw the human view from the saved data, not from the model's text, and limit the "live" additions (spinner, duration) to what is really known at that moment.
Chapter 6 · commit 6c47266 · journal.ts, recovery.ts, deepseek.ts

The journal and recovery: interrupted work must not be repeated blindly

The symptom. The app can be closed, the connection can drop, or the user can open an old conversation tomorrow morning. Without a record, the agent might propose again a change that has already been applied, or conversely think a change that actually failed has succeeded.

interface ActionEvent {
  id, question, tool, time: string
  status: 'proposed' | 'applied' | 'rejected' | 'failed'
        | 'interrupted' | 'read' | 'reverted'
  changes: Change[]       // before/after snapshots → the diff can be reopened
  summary: string
}

Every tool call is recorded. A proposal is recorded as proposed before the review window opens, and its status is updated after the decision. If the process dies between the two points, the journal still holds proposed. On the next turn, reconcileEvents() compares it with the actual contents:

Simulation 5: reconciling an interrupted proposal Interactive · real logic

A batch of three changes had the status proposed when the app was closed. Choose the state of each file as found when the conversation is opened again.

Journal lines sent to the model (as data, not instructions)

The journal and the work status are inserted as a message labeled "data, not instructions", together with an order not to repeat actions that have been applied or rejected. When loaded from disk, parseEvents() only accepts entries with a valid shape (a clean path, a known kind). That metadata can be edited by the user, so its contents are only restored as data and never executed.

A retry that does not duplicate text

// recovery.ts: retry only if no output has been shown yet
if (!(e instanceof TemporaryProviderError) || emitted || attempt >= 2) throw e
await wait(attempt === 0 ? 300 : 900)

Temporary errors (429, 5xx, a failed connection, or a cut-off stream) are retried at most twice, but only if not a single piece of text or reasoning has been shown yet. Retrying after part of an answer has streamed in would duplicate text on the user's screen.

A history summary that does not make things up

When the history exceeds 25% of the budget, the oldest turns are cut off in pairs. The cut-off part is replaced by an extractive excerpt (≤400 characters per turn, ≤4,000 in total), not a model-written summary. Excerpts add no new facts and need no extra API call. The full history stays on disk.

LessonRecord the intent before acting and record the result afterwards. When recovering, reconcile with the real world, not with memory. There are three possible outcomes: it already happened, it has not happened, or unclear, check first. The third possibility must always exist.
Chapter 7 · commit 1c90735 · src/agent/trace.ts, ui/logviewer.ts

The Agent log: seeing what the agent actually did

The need. When the agent answers strangely, the question is almost always the same: what was sent, which tools were called with which arguments, and what came back? The step cards in the panel only show a one-line summary.

class AgentTrace {
  events: TraceEvent[]        // ≤ 2,000, the oldest dropped
  begin(kind, key, title)     // status 'running'
  append(kind, round, delta)  // text/reasoning streaming in
  finish(kind, key, status, detail)  // + duration (ms)
}
TraceKind = 'turn' | 'round' | 'reasoning' | 'text'
          | 'tool' | 'usage' | 'note' | 'error'

Every model round records the number of messages, the tools offered, and the tokens (input / from cache / output). Every tool records its tidied JSON arguments and the exact result returned to the model. Details are cut off at 20,000 characters.

The log is not sent to the model and is not part of the conversation's Markdown file. The assistant's log is saved beside it as <name>.log.jsonl and restored when the conversation is reopened; a New conversation starts an empty one. The pi log only exists in memory.

LessonObservability for developers and context for the model are two different channels. If the log is sent to the model too, the model will "read itself" and the tokens will balloon. Separate the two from the start, and make sure the log is bounded.
Chapter 8 · commit 90c28bb · changes.ts, batch.ts, gittools.ts, markdown/lint.ts

Broader actions, with limits just as tight

With the foundation of Change, preflight, the journal, and verification, new capabilities can be added without opening a shortcut. Every new tool uses the same approval path.

CapabilitySafeguard
insert_text at the start, the end, or after a lineFor after_line, the model must include line_text. If the number and the contents do not match, the proposal is rejected.
edit_file with all=trueMust be stated explicitly. Without it, old_text must still be unique.
delete_fileThe file is moved to the Trash, not deleted permanently. The snapshot in the journal makes Undo possible.
move_fileThe destination must not already exist. The path is checked for symlinks with projectPath().
Git history (git_log, show_commit, file_at_commit)Read-only. The pathspec is limited to Markdown files that are not hidden. A commit may only be a hash or HEAD~n. Results ≤ 6,000 tokens.

8a. Partial approval: why one Change per file matters

The batch review window now has a checkbox per file. Partial approval is only safe if each Change stands on its own. That is why planBatch (Chapter 5) already merged actions per file, and why files that are deleted or moved are locked so that no other action in the same batch may touch them. When the user applies only part of a batch, the journal records two events: one applied and one rejected. The model is told that the rejected part must not be repeated without asking. The Note for the agent box is passed along with the decision (≤2,000 characters), so a rejection can come with a reason.

8b. Undo = the inverse, passing through the same preflight

invertChange(c):  create → delete    delete → create (from the before snapshot)
                  edit   → edit with before/after swapped
                  move   → move from c.to back to c.file
// applied as a reversed batch: [...changes].reverse().map(invertChange)

Because the inverse is also a Change, Undo automatically fails safely if the file has been edited again. You already tried it in Simulation 4. The journal marks it reverted so that the agent does not think the change is still in effect.

8c. Structure verification: only new problems cause a failure

The structure check (markdown/lint.ts) looks for tables whose cell counts do not match, unclosed code blocks, empty headings or ones that skip a level, and links to Markdown files that do not exist. Real documents often already have such problems, so the check is compared with the contents before the work (the baseline, taken from the before of the first change per file). Every problem has a key without a line number, so an old problem that only shifted lines is not counted as new. Only problems that appear because of the agent's change make it fail. The criterion file: "*" checks the whole folder, for example to make sure an old date remains nowhere.

The review window for a kanban board change proposal
Kanban board changes also go through the same review window. Board serialization may tidy the spacing elsewhere, and that is shown as it is in the diff.

8d. A tool must always answer: the Git result that never came

The agent loop waits for every tool call before the next round. A read tool that never returns does not fail; it freezes the whole turn, with the spinner still turning. That is what an occasional timeout of the test agentGit in a real repository turned out to be. The log of a failing run showed:

Gjs-CRITICAL: Attempting to run a JS callback during garbage collection … it has been blocked.
The offending callback was AsyncReadyCallback().

runGit() (git.ts) starts git with Gio.Subprocess and waits for communicate_utf8_async(). Git had exited (GLib had already reaped it) and timers still ran, but GJS 1.80 blocked the completion callback because it arrived during a garbage collection, so the Promise behind git_log never settled. This is the same runtime guard described in bench/GC-DIAGNOSIS.md; it is rare (about a third of the test runs with every CPU core busy, and now and then without load) and cannot be fixed from the app.

runGit(cwd, args):  start git, wait for communicate_utf8_async()
  watchdog every 500 ms:
    git has exited and no result after 5 s   → failure "git finished but its result was lost; try again"
    git still running after 120 s             → stop git, failure "git did not finish within 120 seconds"
  the first of (callback, watchdog) settles the run; the other is ignored

Both endings are ordinary failures: formatGit() turns them into a message the model reads, and it can call the tool again. The run is not repeated automatically, because the same runGit() also commits for the History tab, and a commit whose result was lost may already have happened. The second limit covers a git that waits for input that will never come, such as a signing passphrase. The test replaces Gio.Subprocess.prototype.communicate_utf8_async with a function that never calls back, exactly what the runtime does, and also runs a git alias that sleeps longer than the limit; with the watchdog disabled, the test times out.

LessonEvery wait in an agent loop needs an end that does not depend on the thing being waited for. A failure the model can read is recoverable; a tool that never answers is not.
LessonInvesting in the right structure early (changes as data, one change per file, preflight) makes later features such as partial approval and Undo small and safe. A new feature does not need to make its own write path.
Chapter 9 · commit b1ab712 · agent/harness.ts, orchestrator.ts

Markwork as an orchestrator: another agent works, Markwork holds the board

The need. Project code (for example ~/web-ecommerce) does not live in Markwork's work folder, and the work needs an agent that can run commands, edit code, and run tests. That is not a job for the Markwork agent, which deliberately never writes by itself. The answer: let another agent (pi) do the work, and make the Markwork kanban board the control plane. The card states the task, @pi states who does it, and the list states its status.

## Plan
- [ ] Checkout with QRIS @pi #feature
  Use the official SDK.
- [ ] Test the cart @pi #project/shop-admin

// board frontmatter: project: web-ecommerce
// settings.json: projects = {
//   "web-ecommerce": "/home/eka/web-ecommerce",
//   "shop-admin":    "/home/eka/shop-admin" }
interface Run {
  board, card        // board + card text (the identifier)
  agent: 'pi'
  project, folder    // name → folder from the settings
  prompt
  status: 'queued' | 'working' | 'done'
        | 'failed' | 'stopped'
  trace: AgentTrace  // the pi log
  result: HarnessResult | null
}
right-click → Work on it with pi→ buildPrompt + RunQueue→ pi --mode json (cwd = project folder)→ JSONL PiReader → AgentTrace→ card: In Progress → Review + ↳ summary

9a. The trust boundary moves to the folder

Up to Chapter 8, every write went through Change and the review window. An external harness cannot be treated like that: pi writes directly, runs bash, and uses its own permission rules. So the boundary is no longer per change, but per folder:

RiskSafeguard
A card, a note, or the Markwork agent points the harness at an arbitrary folderMarkdown only mentions the project name. The name is mapped to a folder in settings.projects, which is only filled in by the user through the folder chooser dialog.
The harness writes into Markwork's work folder without a review windowcheckProjectFolder() rejects the work folder, its contents, its parents, the root, and relative paths.
Two harnesses overwrite one repoRunQueue: one working run per folder, the rest queued in order.
The harness waits for input and nobody knowsA permission/input request becomes a ⏸ waiting status on the card and is answered from Markwork (9d). After a turn finishes without a question, stdin is closed so that pi exits.
A grandchild process holds the pipe, the run never finishesAfter the process exits, the remaining output gets a 1.5-second grace period and is then abandoned. Stop = SIGTERM, then SIGKILL after 5 seconds. Closing the window stops all runs.
Changes come in without being reviewedThe prompt asks pi not to commit/push. The card stops at Review, and the diff is reviewed in the project repo.

9b. Delegation through a process contract, not a tool call

pi is not called as a model tool. It is a separate process with a clear contract: arguments in, JSONL out, an exit code at the end. PiReader reads the events tool_execution_start/end, message_update, message_end (with stopReason and usage.cost), and agent_settled, and maps them to the same AgentTrace as in Chapter 7. The Agent log window is reused without any change. Lines that are not JSON are recorded as they are and do not make the reading fail.

One GJS trap: the pipe from Gio.Subprocess.get_stdout_pipe() is wrapped as a Gio.UnixInputStream after Gtk is loaded, and GJS prints a Gjs-WARNING. That is why the pipe is read through GLib.IOChannel, split only at LF according to pi's JSONL framing.

9c. The board as the source of truth, the card looked up again

The board has no card ids, so a card is recognized by its text. Every time the status changes, locateCard() looks for it on the current board (the user may have moved it) and refuses if it is missing or not unique. A board that is open is changed through its editor (one undo step); a closed one is written straight to disk.

Simulation 6: the queue per project folder Interactive · real logic

Assign several cards to two projects, and then finish, fail, or cancel their runs. What runs is a copy of RunQueue.add() and end() from agent/harness.ts.

No cards have been assigned yet.

9d. Waiting for an answer: from a one-way stream to a two-way conversation

The symptom. The first version used pi --mode json with stdin closed right away. If pi asked ("Do you want SDK A or B?"), the card still went into Review and the question was simply left behind in the note. If a pi extension asked for permission (for example before rm -rf), the dialog had no recipient: the permission-gate shipped with pi's examples actually blocks the command when there is no UI.

The design. pi is started with --mode rpc. Its process stays alive for the run, receives commands through stdin, and sends the same events as the JSON mode, plus two new things:

// stdout → Markwork
{"type":"extension_ui_request","id":"u1",
 "method":"confirm","title":"Allow the bash command?",
 "message":"rm -rf build"}
{"type":"agent_settled"}    // the turn is finished

// Markwork → stdin
{"type":"extension_ui_response","id":"u1","confirmed":true}
{"id":"markwork-2","type":"prompt","message":"Use SDK A"}
{"id":"markwork-steer","type":"steer","message":"…"}
PiSignal = { type: 'ask', ask: HarnessAsk }
         | { type: 'settled' }
         | { type: 'rejected', error }

HarnessAsk {
  kind: 'select' | 'confirm' | 'input'
      | 'editor' | 'question'
  id     // null for a question
  title, message, options, prefill
  timeout  // pi answers by itself after it
}
PiReader.line() → ask→ run.status = 'waiting' · chip ⏸→ Answer pi… (Allow / Deny / text / Later)→ extension_ui_response / prompt through stdin

There are two kinds of "waiting" with different answer channels. An extension dialog has an id and is answered with extension_ui_response. A model question has no special protocol: pi just stops with some text. Markwork treats it as a question if the last answer ends with ?, and the reply is sent as the next prompt to the same process. The reader discards the old turn state through restart() so that an error or answer from before is not carried over. This heuristic can miss, so every finished run can still be replied to through Reply to pi…, which runs pi again with --session <id> and the same log.

Simulation 7: is pi asking a question? Interactive · real logic

Write pi's last answer. A copy of endsWithQuestion() decides whether the card waits for your reply or stdin is closed so that the card goes into Review.

LimitReason
A waiting run still holds its folder's queueIts process is still alive and can continue the work as soon as it is answered; another harness in the same folder could collide with it.
Later closes the dialog without answeringAnswering is a decision. Closing a dialog must not be read as "allow" or "deny".
An extension's timeout is followedpi uses its default answer after the time limit. The status on the card goes back to working too, so that it does not show a request that has gone stale.
One command = one JSON line, bytes with an explicit lengthpi's RPC framing is split on LF. write_chars() with a string and length -1 in GJS triggers GLib-WARNING: Invalid UTF-8, and the steering that is sent is corrupted.
Measured with a real pi (once): a card that asked it to ask a question first went into the waiting status with the question "What is the note file's name?". After it was answered, a scratch extension asked for permission for ls, permission was given from a script, and then pi created the file and the card went into Review (8,699 tokens, ±$0.0037). That is a single run, not an evaluation. Your own pi does not have a permission extension yet, so without installing one only questions will appear. The custom() dialog (TUI only) is not supported by RPC mode, and there is no interactive terminal yet.
LessonGood delegation still needs a return channel. Tell apart requests that have a protocol (an id + a structured answer) from those that are only implied in text (a model question). The first is answered exactly. The second is guessed simply, and then give a way out when the guess misses (Reply).

9f. Notes as context: the contents are sent, the access is not

The need. Design decisions, specifications, and meeting minutes live in Markwork's work folder, while pi works in the project repo. The boundary in 9a deliberately makes pi unable to read the work folder. As a result, the card "Checkout with QRIS" only carries a single title line. Users had to copy the specification into the card note by hand, and that copy went stale quickly.

The design. A card links notes with the syntax already used in the editor, [[Specification#Payment]]. On Work on it with pi, Markwork resolves the link and copies its contents into the prompt. Pi still gets no path to the work folder, and the boundary in 9a does not change.

// pure: agent/harness.ts + markdown/wikilink.ts
cardWikiLinks(card)        // title + note, ≤ 10,
                           // no duplicates, code skipped
LinkedNote {
  link: WikiLink           // target, heading, alias
  file: string | null      // null = not found
  text: string | null      // null = #section is missing
}
buildPrompt(card, project, board, notes)
// host (window.ts), same rules as Ctrl+click
linkedNotes(boardFile, links):
  root  = work folder (or the board's folder)
  files = workspace.names(root)    // no dot folders,
                                   // no symlinks
  file  = resolveWikiLink(target, files, board)
  text  = open editor ?? disk
  text  = heading ? noteSection(text, heading) : text
Work on it with pi→ cardWikiLinks→ host.linkedNotes (reads the work folder)→ linkedNotesSection: fence + limits→ prompt (shown in full in the pi log)
RiskSafeguard
A link is used to read files outside the work folderThe target is only matched against the workspace.names() list: Markdown files in the work folder, without dot folders and without following symlinks. No path from Markdown is opened directly.
A long note inflates the prompt and pi's costAt most 10 links, 8,000 characters per note, and 24,000 in total. Whatever is cut is marked with the number of characters dropped. #section copies only a single section.
A note's contents hold ``` and "escape" the block, being read as instructionsThe block fence is made one backtick longer than the longest run of backticks in the contents. The prompt states that the contents are context, not commands.
The user does not know what was sentThe full prompt, including the note contents, is recorded in the pi log. A link that is not found is mentioned as it is, not silently dropped.

Simulation 8: related notes in the prompt Interactive · real logic

Change the card or the contents of the "Specification" note. What runs is a copy of wikiLinksIn(), noteSection(), and linkedNotesSection(). The sample work folder holds specs/Specification.md and Meeting Notes.md. The limit in the simulation is reduced to 300 characters per note so that the truncation is visible.


  
What has been measured and what has not: 4 unit tests (card links, a prompt with a section/missing/truncated note, the total limit) and 1 GUI test with a fake pi check that the prompt reaching the process contains the #Color section but not the other sections, that a missing note is still mentioned, and that clicking a link on the card does not also open the edit dialog. It has not been run with a real pi. It has not been measured whether the model uses that context better than a manual copy. The contents are taken when the run starts. Reply to pi does not copy them again, so a note changed midway is not sent.
LessonWhen a delegated agent needs data from territory that is deliberately closed to it, do not open the access. Send the contents: the host resolves the reference with the same rules as for the user, limits its size, wraps it so it is not read as a command, and records exactly what was sent.

9e. Testing: a real process, a fake harness

The GUI tests run a shell script that pretends to be pi, through the same process path. These tests check the cwd and the prompt, a card moving to In Progress and then Review with a note, a queue that is truly in order (the script records start/end), Stop on a process stuck for 30 seconds, a failure with a message from stderr, a project folder that is asked for once and then remembered, and a closed board that is updated on disk. For RPC mode, the fake script reads get_state and prompt from stdin, sends a permission request or a question, and then records the answer it received. The tests check the exact answer that arrives ({"type":"extension_ui_response","id":"ui-1","confirmed":true}), Later still waiting, End without replying, steering that arrives as steer, and a reply that runs with --session test-session in the same log. The unit tests cover @name assignment (not an email or @{date}), the project, folder rejection, the prompt, PiReader, the result note, and the queue. There are 23 tests, all without a real pi or API.

What has been measured and what has not: one end-to-end test with a real pi 1.0.3 (DeepSeek, the task "create HALO.txt" in a sample repo) succeeded. The file was created in the project repo without a commit, the work folder stayed empty, the card moved to Review, and the run used 4,248 tokens (±$0.0038). That is a single run, not an evaluation. Parsing @name adds ±0.35 ms per 500 cards in cardMeta() (±0.7 µs per card). Not there yet: the waiting status when the harness asks for permission, an interactive terminal, a worktree per card (cards in the same project queue up), and a diff view of the harness result in Markwork. The run status only exists while the app is open; what persists is the card position and the ↳ note.
LessonWhen delegating to another agent, do not pretend to be able to review every step. Move the boundary to a place that can be guarded: where it may work (a name → a folder chosen by the user), how many at a time (a queue per folder), and when to stop (EOF, grace period, SIGKILL). Then make the result visible where the user already works: the card, the note, and the log.
Chapter 10 · commit 1bac3b8 · agent/commitmessage.ts, ui/messagewriter.ts

The smallest model call: one line for the commit box

The need. The History tab can commit several files at once, but the message is still typed by hand, and a writer's commits easily turn into "update plan.md". The model can see what changed, so it can suggest a better line. This needs no conversation, no tools, and no proposal. It is one request whose answer goes into a box the user already reviews.

What is sent

Only the diffs of the checked files against HEAD (workingDiff) and the subjects of the last 10 commits (recentSubjects), so the message follows their language and style. No tools are offered, and Deep thinking is off: the answer is one line, so speed matters more than reasoning.

Where the answer goes

Into the commit message box, streamed and cleaned with cleanMessage() (first line, no code fence, label, quotes, or trailing period). The box is selected afterwards, and Ctrl+Z restores the previous text. The commit itself still needs the Commit button.

// agent/commitmessage.ts: no GTK; git and the key come in as CommitSources
writeCommitMessage(sources, files, onText, cancellable)
  key   = sources.keyStore.get()            → none: { ok: false, reason: 'no-key' }
  diffs = files.map(sources.diff)           → one unreadable diff is named, not fatal
  fit   = fitDiffs(diffs, 30_000 chars)     → small diffs whole, the largest cut to whole lines
  chat({ messages: commitPrompt(fit, subjects), thinking: false, onText })
  return { ok: true, message: cleanMessage(answer), shortened, cancelled }

The budget. A diff can be as long as a book. fitDiffs() shares a fixed budget out fairly: going from the smallest diff up, each gets at most an equal share of what is left, so small diffs are sent whole and only the largest ones are cut. The cut ends on a whole line and says how many lines were left out, and the prompt tells the model that some diffs were shortened. The status line under the box tells the user the same thing.

Simulation 10: sharing the diff budget Interactive · real logic

Three checked files with diffs of different sizes. What runs is a copy of fitDiffs(). The dark part of each bar is sent; the light part is cut.

The boundary. The button reads only what the user is about to commit, which is the same set the user already chose with the checkboxes. While it writes, MessageWriter locks the box, the checkboxes, and Commit, so the files and the message cannot change under it. The button becomes Stop, and Stop keeps the text written so far. A missing key, a network error, or an empty answer puts the previous text back and says why in the status line. Without a key, the status line links to the Assistant settings. The window gives the button its writer (connectWriter) using the Assistant's key store, provider, and model, so tests swap in a fake writer without touching the network.

One GTK detail. Gtk.TextBuffer.set_text() is not recorded in the box's undo history, so GTK's own Ctrl+Z cannot step back to the user's draft. MessageWriter keeps that one step itself: a capture-phase key handler restores the previous text while the box still holds exactly the written message, and forgets it as soon as the user edits.

What has been measured and what has not: 10 unit tests (the budget and its line-aligned cut, the cleanup, names relative to the common folder, the prompt, a fake provider that streams, no key, one or all diffs unreadable, a provider error) and 2 GUI tests with a fake writer (streaming into the box, the lock, Stop, Ctrl+Z, the failure and no-key states, the disabled button without checked files, and the same button in the single-file window). It has not been run against the real DeepSeek API, so the quality of the messages and how well they follow the style of earlier commits are not measured. The 30,000-character budget is a judgment, not a measurement.
LessonNot every model feature needs the agent loop. When the input is already chosen by the user and the output lands somewhere the user reviews anyway, the safest shape is the smallest one: a single request without tools, a hard budget on what is sent, and an answer that only fills a field. The user stays the one who acts.
Cost · bench/AGENTIC.md · npm run bench:agentic

What does that safety cost?

Snapshots, verification, and the journal are not free. The following measurements run locally without calling a model: 10 repetitions after 3 warm-ups, a batch of 20 files, at sizes of 30,000 and 480,000 characters per file. The fixture replaces the entire contents of a file, so it is deliberately heavier than ordinary work.

Operation (20 files)30,000 chars: medianp95480,000 chars: medianp95
Plan the batch8.98 ms12.00 ms113.94 ms122.10 ms
Verify the result0.25 ms0.33 ms3.43 ms4.73 ms
Checkpoint without the journal3.52 ms4.51 ms3.46 ms4.61 ms
Checkpoint with the journal19.46 ms31.58 ms188.56 ms212.64 ms
Structure verification (with a baseline)10.11 ms10.81 ms157.35 ms160.57 ms

Optimizing the new code in the same piece of work (90c28bb)

The first version of the structure check walked every document three times (frontmatter/code, tables, links) and trimmed whitespace on every line. The "*" check split all files into lines even if the text being looked for was not in that file. Now the three are merged into a single pass with a character filter before the regex, and "*" only splits files that actually contain the text (file.text.includes() first).

First versionAfter merging (90c28bb)
Honestly unfinished: a checkpoint with the journal at 20 × 480,000 characters takes ±190 ms on the main thread and is still written synchronously. The 157 ms structure verification is also felt as a pause, although it only happens once per verify_work call, not while typing. Journal snapshots can make the conversation file large. This benchmark does not measure API latency, how accurately the model picks criteria, or how long the user takes to review a diff. An evaluation of a real model is available through npm run test:live -- --agentic (it uses API quota) and has not been run in this measurement session. The editor benchmark (bench:compare) shows no repeated regression.
Summary

Patterns that carry over to other agents

PatternStructureIn Markwork
Context split by how often it changesstable system + changing notebuildContext, buildMessages
A loop with hard limitsMAX_ROUNDS, toolBudget, the last round without toolsChatSession.ask
Pure tools that do not throw(args, SourceFile[]) → { content, summary }runTool, formatGit
Changes as dataChange { before, after }planChange, onProposal
Re-check right before writingpreflight(c, read)applyBatch, Undo
Plan on a virtual copya copy of SourceFile[] + locksplanBatch
Done = passes the checksWorkState.status + VerificationverifyWork, structureCheck
One state, two readersstored data → text for the model, a widget for humansworkText, WorkList
Record the intent before actingActionEvent, three-outcome reconciliationreconcileEvents
Observability separate from contexta bounded in-memory logAgentTrace
Delegation through a process contractarguments → JSONL → exit code, mapped to the same logPiReader, spawnHarness
Names, not paths, for dangerous placesa name in the data, the mapping only from the usersettings.projects, checkProjectFolder
A return channel for the delegated agentrequests with an id are answered exactly; text questions are guessed + a way outHarnessAsk, answer, resume
Context sent as contents, not accesslinks in the data are resolved by the host, the contents copied with limitslinkedNotes, linkedNotesSection
One writer per resourcea queue per key (folder)RunQueue
The smallest call that does the jobone request without tools, a hard input budget, the answer only fills a field the user reviewswriteCommitMessage, fitDiffs, MessageWriter
Every wait has its own enda watchdog that does not depend on the awaited callback; the first to finish winsrunGit, gitLimits
Checklist before adding an agent capability 1) Does the tool read or change? A tool that changes things must go into CHANGE_TOOLS and only produce a Change. 2) What is its limit? Decide the path, the result size, the count, and what happens when the limit is exceeded. 3) What does the user see before approving, and how can it be undone? 4) How is the result verified, and how does this capability recover after an interruption? 5) Write tests with a fake Provider, measure the cost with bench:agentic, and then update this page.

Sources: README.md (the "How the assistant works" section), AGENTIC.md, src/agent/, and the git commit history from f6a4cc7 to 90c28bb, plus the harness orchestrator (Chapter 9). See also learning performance from the Markwork editor.