Unroll mhermon.github.io/unroll/guide

Using Unroll

This guide covers opening a file, reading a run, following tool calls, checking steps against the source, searching, selecting runs in files that hold many of them, and the optional investigations. To follow along, download the example dataset of agent runs or the example Claude Code session.

Open a file

Right-click any .jsonl, .ndjson or .json file and choose Unroll: Open as Agent Trace. The same command is on the editor’s preview icon and in the Command Palette, and Reopen Editor With… → Agent Trace (Unroll) works too.

Your usual text editor stays the default for these files, so nothing changes until you ask for Unroll.

If your results are on another machine, open that folder over Remote SSH, WSL or a dev container. Unroll runs there and reads the file in place, so large results don’t need to be copied.

Read a trace

Each step shows its role, the line it came from in the file (L4) and when it happened. Times are shown in UTC; hover one for the full timestamp.

  • Reasoning is folded by default, with its length shown. Select it to read it.
  • Tool calls show the tool name and its arguments as readable fields. Expand a call to see every argument.
  • Long messages start as an excerpt. Select Show all to read the rest; when you search, the matching passage is shown.
  • Collapse any step with Enter or by selecting its label. The View menu collapses or expands all loaded steps and switches between comfortable and compact spacing.

Follow a tool call

When a call’s result is in the file, the call shows result at line N. Select it to see the call and everything it returned together, even when the result arrived many steps later. Select Clear or press Esc to go back to the whole trace.

Calls and results are matched by the call IDs in the log. If a match is ambiguous, Unroll leaves the call unlinked rather than guessing.

Check the raw data

Hover a step to show its actions:

  • Record (R) switches the step to the exact JSON record it came from. In a dataset, you can switch to the whole row.
  • Ln (O) opens the file beside the preview at that line.
  • Copy copies the step, or the record when it’s shown.

Records Unroll doesn’t recognize, and lines that aren’t valid JSON, are hidden from the conversation but never dropped. Turn on View → Show events to see them. Very large records show a labelled preview; open the source to read all of it.

Press CtrlF (CmdF on macOS) or / to search message text, tool names, arguments and results. Matches are highlighted, including inside formatted arguments.

The role tabs (All, User, Assistant, Tool, System) filter the conversation. Line numbers always refer to the original file, so a filtered view still points at the right place.

On a large file, search keeps reading through the file and shows Stop while it works; Continue reading picks up where it stopped. Esc clears search and filters and returns you to where you were reading.

Datasets of many runs

Eval results and training datasets often hold one full conversation or agent run per row. Unroll reads these as a list of runs. It recognizes SWE-agent, OpenHands, τ-bench, AgentInstruct, Hermes and xLAM rows, ShareGPT-style chats, and JSON files that wrap their rows in one key, such as {"data": [...]}.

  • Run list. Each row shows its id and, when the row has one, its outcome (resolved, success, reward, exit_status and similar). Press C to show or hide the list.
  • Filter. The box above the list matches ids, row fields and message text. Rows that match only inside their messages are marked in text, and opening one highlights the match. To narrow by a field’s value, add a condition.
  • Row fields. The row’s other fields, such as model, reward, cost or a generated patch, appear above its conversation. Long values are folded.
  • Moving between runs. ⇧J and ⇧K open the next and previous row. In the list, use ↑↓ and Home/End, or type a number in the Row box.

Select cases

Conditions pick the runs worth reading, such as the failures a judge scored highly. Select + Add condition above the run list, choose a field, and then how to compare it.

  • Fields. The list shows the fields in the runs you’ve loaded, with nested metadata under its parent, and what each holds: text, number, true/false or a list. Any field can be typed as a path, such as metadata.judge.score. Quote a key that contains dots: scores["pass@1"].
  • Conditions. Only the comparisons that fit the field are offered: is and is not with one or more values, contains, > ≥ < ≤ for numbers, is true, is null, exists and is missing. Is not also matches runs without the field. Next to text comparisons, Aa matches case and .* uses an RE2 regular expression.
  • Values. Suggestions come from the whole file, most common first, with how many runs hold each. Large files are counted from an even sample, and the list says so. Number fields show their range as you type.
  • Lists. A condition on a list matches when any item does, and reads any tags is "math". Its opposite reads no tags is "math".
  • Combining. Match all or any of the conditions. Pause one with its checkbox, or select it to edit. ⇧J and ⇧K move between matching runs only.

A run whose field is too large or holds an object can’t be compared; the list says how many weren’t checked rather than leaving them out silently.

Watch a run live

Leave a trace open while an agent writes to it. New steps appear as lines are added, and unsaved edits show up too. Turn on Follow latest (F) to keep the newest steps in view; scrolling up or changing a filter pauses it.

A half-written last line doesn’t hide the steps before it.

Large files

Saved .jsonl and .ndjson files have no size limit. Unroll reads the part you’re looking at straight from disk, so a 100 MB log opens about as quickly as a small one.

  • Scroll to load earlier or later messages, or press [ and ]. The buttons in the footer jump to the first or last messages.
  • The map on the left (T) shows the loaded messages, colored by role. Select a mark to jump to it.
  • While a new file is still being indexed, totals are marked +.
  • .json files, unsaved edits and files on virtual filesystems are read into memory instead, up to 100 MB by default. Change this with the Unroll: Max File Size MB setting.

Size of what you’re reading

The In view total in the header estimates the size of the loaded messages. Select it to switch between estimated tokens, words and characters. Turn on per-message counts in View.

Counts cover the messages, reasoning and tool activity currently loaded, so they change as you scroll or filter. They’re estimates of what you’re looking at, not a total for the whole file, the size of a model request, or what you were billed.

Investigations (experimental)

Investigations let you ask a language model about a run or a dataset, such as why a run failed or how often agents edit tests before running them. The model reads the file through read-only tools and cites the steps it relied on. For questions about how often something happens, it asks you to label a random sample of runs; the estimate and its 95% interval come from your labels, not the model’s. Investigations are off by default, and nothing else in Unroll needs them.

  • Turn them on with Unroll: Enable Investigations (Experimental)…, which explains what is sent before anything changes. Unroll: Turn Off Investigations hides them again.
  • Choose a model with Unroll: Choose Investigation Model…: a VS Code language model such as GitHub Copilot, the Anthropic API with your key, or an API endpoint (its base URL, Responses or Chat Completions, a model ID and that service’s key). Keys are kept in VS Code’s secret storage and used only for the endpoint they were entered for. An endpoint on another machine over plain HTTP shows a warning, because excerpts would travel unencrypted; use HTTPS where you can.
  • Ask: add steps with X, Ctrl-click or a text selection, then press A. I shows or hides the investigation panel.

When you ask, Unroll sends your question, the steps you selected with their context, and what the model reads with its tools, all from the open file. Before the first request to a provider and model, you see where excerpts will go and confirm. Remote providers bill you; the panel shows the steps used and the cost or tokens so far.

  • Investigations run only in trusted workspaces, and their settings can be changed only in your user settings, so a repository you open can’t turn them on or send them elsewhere.
  • The model’s tools only read the file. Unroll never runs recorded tool calls, whatever a trace or model says.
  • Results are saved in Unroll’s private storage. Set Unroll › Investigate: Save To to workspace to keep them in .unroll/investigations. Saved files include trace excerpts and the file’s path, so review them before committing.
  • Unroll: Remove Stored Model Keys deletes every stored key and confirmation.

Keyboard

Press ? in the preview for the full list.

KeysAction
JKNext / previous step
NPNext / previous tool call
⇧J⇧KNext / previous run in a dataset
RShow the original record
OOpen the step’s line in the file
/ or CtrlFSearch
EscClear search and filters
FFollow the newest steps
[]Load earlier / later messages
CShow or hide the run list
TShow or hide the map
EnterCollapse or expand the focused step

Single-letter shortcuts are ignored while you type in a field. Every shortcut is a VS Code command named Unroll: …, so you can rebind it under Keyboard Shortcuts or run it from the Command Palette.

Formats

Files can be JSONL, NDJSON or a single JSON array. Unroll recognizes these record shapes:

SourceRecognized content
Claude Code, AnthropicMessage envelopes, text, thinking, tool use and tool results
Codex, OpenAI Responsesresponse_item and event_msg records, messages, function calls, custom tools and outputs
OpenAI Chat Completionsrole / content messages and tool_calls
GeminiModel messages, parts, function calls and responses
Vercel AI SDKText, reasoning and tool parts
OpenTelemetry GenAIgen_ai.input.messages and gen_ai.output.messages span attributes
LangChain, LangGraphHuman, AI and tool messages, including nested state dumps
ShareGPT, chat datasetsConversation and message envelopes
Trace datasetsOne run per row or array item: SWE-agent trajectory, OpenHands messages, τ-bench traj, AgentInstruct and Hermes conversations, xLAM-style query / answers, and thought + action turns with an environment reply (Agent-SafetyBench, ReAct-style logs)

Log formats change between tool versions, so a file may only partly match. Anything unrecognized is still available with View → Show events. Parquet isn’t supported. Provider and benchmark names describe compatible formats, not affiliation.

Troubleshooting

  • No messages appear. Turn on View → Show events to see what the file contains, and compare it with the source. If it’s a format you expected to work, please report it.
  • The preview doesn’t load. Close and reopen it. If that doesn’t help, check the Unroll output channel for errors.
  • A file is too large to preview. This only applies to .json files, unsaved edits and virtual filesystems. Raise Unroll: Max File Size MB, or save the data as .jsonl.

Report a problem with your extension version, VS Code version, OS and a small sanitized example. Issues are public, so never attach credentials or private conversations.

Requirements

  • Desktop VS Code 1.96 or newer on Windows, macOS or Linux. The browser-only vscode.dev isn’t supported.
  • Remote workspaces (SSH, containers, WSL) work. Unroll runs wherever VS Code runs its extensions, so the file is read on that machine.
  • Settings and command IDs start with chatview. for compatibility with earlier builds.

No telemetry or account. Reading never contacts a model provider; only investigations you turn on do. Privacy details