Unroll mhermon.github.io/unroll
unroll.jsonl 20 records · read-only In view ≈0 tokens
GitHub
1
20
userL1

Our agent failed 58 of 200 tasks in last night’s eval. The results file is 110 MB of JSONL. How do I find out what went wrong?

assistantL2
reasoning31 words

They need to read the failed runs, be able to trust that what they see matches the file, and skip writing another parsing script. Start with the runs, then the steps.

Read what your agent did,
step by step.

Unroll is a free VS Code extension for reading agent transcripts: eval runs, Hugging Face trajectory datasets, and Claude Code or Codex sessions. It shows each run as a conversation, links every step to the JSON record it came from, and lets you pick the runs worth reading by outcome or any other field. Files are read locally and nothing is sent anywhere.

Install for VS Code Try it on this page

Unroll runs in desktop VS Code. Send yourself the link and install it there.

Free and MIT licensed · VS Code 1.96+ · Windows, macOS, Linux · works offline

Unroll { "file": "claude-session.jsonl" } → result at line 3
toolL3
↳ Unroll · line 2

This is a Claude Code session, 13 lines of JSONL as the agent wrote them. Press play or drag the handle to show it as a conversation, then select any message to see the line it came from. The other tabs have a dataset of agent runs, a chat dataset, other log formats and a 50,000-line file. You can also open one of your own files; it’s read in this tab and not uploaded.

raw unrolled 0 / 13

    Drop a .jsonl or .json file anywhere on this box. It is parsed in this tab and never uploaded.

    assistantL4

    That was one session. An eval usually produces one file with hundreds of runs, so the first job is choosing which runs to read.

    Select { "where": "resolved is false" } → result at line 5
    toolL5
    ↳ Select · line 4

    Pick the runs to read

    When each row of a file is a full run, as in SWE-agent, OpenHands and τ-bench results, Unroll lists the runs with their outcomes. Add conditions on any field, including nested metadata such as metadata.judge.score, to narrow the list to the runs you care about. Each run’s model, score or cost appears above its conversation, and ⇧J and ⇧K step through the matching runs.

    A dataset of eight agent runs narrowed by the condition resolved is false to the three failed runs, with the most expensive one’s model, cost and conversation A dataset of eight agent runs narrowed by the condition resolved is false to the three failed runs, with the most expensive one’s model, cost and conversation
    assistantL6

    Inside a run, each tool call is linked to its result.

    Follow { "call": "toolu_01Hx9Rw2" } → result at line 7
    toolL7
    ↳ Follow · line 6

    Follow each tool call to its result

    Arguments are shown as fields instead of escaped JSON strings. Select result at line N to see a call next to its output, even when the output arrives many steps later. Calls and results are matched by their IDs; when a match is ambiguous, Unroll leaves the call unlinked rather than guessing.

    A Bash call filtered by its call ID, shown directly above its failing test output A Bash call filtered by its call ID, shown directly above its failing test output
    assistantL8

    Every step can be checked against the file it came from. Press R on any message on this page to see its record.

    Record { "line": 7 } → result at line 9
    toolL9
    ↳ Record · line 8

    Check every step against the source

    Switch any message to the exact JSON record it came from, or open its line in the file beside the preview. Line numbers always refer to the original file, so a filtered view still points to the right place. Records Unroll doesn’t recognize are kept as raw events rather than dropped.

    A tool call and its result in Unroll on the left, and the same records as raw JSONL on the right with line 4 selected A tool call and its result in Unroll on the left, and the same records as raw JSONL on the right with line 4 selected
    assistantL10

    To find a particular step in a large file, search it. The search box above works on this page too; press /.

    Grep { "pattern": "applyCoupon" } → result at line 11
    toolL11
    ↳ Grep · line 10

    Search the whole file

    Search message text, tool names, arguments and results, and filter by role. Matches are highlighted inside formatted arguments. On a large file, search reads through the whole file in the background, and you can stop it and continue later.

    Searching for applyCoupon highlights the Grep call’s pattern and the matching lines in its result and in the file it read Searching for applyCoupon highlights the Grep call’s pattern and the matching lines in its result and in the file it read
    assistantL12

    A few more details:

    • Large files

      Saved JSONL of any size opens straight from disk, and Unroll reads only what you scroll to. In benchmarks on a laptop, a 110 MB log with 110,000 messages shows its first messages in under a second.

    • Remote machines

      Over Remote SSH, WSL or a dev container, Unroll runs where the files are, so results on a cluster can be read without copying them.

    • Runs in progress

      Leave a file open while an agent writes to it. New steps appear as they’re written; press F to follow the latest.

    • Chat and training data

      Read fine-tuning and preference datasets one conversation at a time, with Markdown, tables and code rendered.

    Unroll also has an optional, experimental way to ask a language model about a run or a dataset. For questions about how often something happens, it asks you to label a random sample and reports an estimate with a 95% interval computed from your labels. Investigations are off by default.

    userL13

    Can I see it working?

    assistantL14

    Here is a one-minute walkthrough, recorded with the extension and made-up example files.

    Play { "video": "walkthrough.mp4", "duration": "1:00" } → result at line 15
    systemL16

    Supported formats

    Unroll reads .jsonl, .ndjson and .json files in these formats. Anything it doesn’t recognize is shown as raw JSON events.

    Claude Code, Anthropic
    Message envelopes, text, thinking, tool use and results
    Codex, OpenAI Responses
    Response items, messages, function calls and outputs
    OpenAI Chat Completions
    Role and content messages with tool calls
    Gemini
    Parts, function calls and function responses
    Vercel AI SDK
    Text, reasoning and tool parts
    OpenTelemetry GenAI
    Input and output message attributes on spans
    LangChain, LangGraph
    Human, AI and tool messages, including state dumps
    Chat and trace datasets
    One conversation per row: ShareGPT, SWE-agent, OpenHands, τ-bench and similar

    Log formats change between tool versions, so a file may only partly match. Inspect eval logs and Parquet files aren’t supported. If your files don’t display well, send a small sanitized sample. Format details

    userL17

    How do I install it?

    assistantL18

    Install

    Unroll works with desktop VS Code 1.96 or newer. Install it from the Visual Studio Marketplace; it takes a few seconds.

    1. Install

      Open Unroll on the Marketplace and choose Install, or search for Unroll in VS Code’s Extensions view.

    2. Open a file

      Right-click a .jsonl, .ndjson or .json file and choose Unroll: Open as Agent Trace. Your usual text editor stays the default.

    3. Offline?

      Download unroll-<version>.vsix from the latest GitHub release, then choose Install from VSIX… in the Extensions view’s … menu.

    Or from a terminal:

    code --install-extension mhermon.unroll
    systemL19

    Your data stays where it is.

    Transcripts often contain unreleased model outputs, proprietary code and credentials. Reading a file in Unroll sends nothing anywhere: there is no account, server or telemetry. Investigations are the only feature that contacts a model provider, and they are off until you turn them on. Privacy details

    • Works with no network connection
    • Reads files on the remote machine over SSH, WSL or containers
    • Never runs the tool calls recorded in a transcript
    • Never changes your trace files
    • Doesn’t load remote images found in transcript content
    assistantL20

    Try it on your own results.

    Install the extension, right-click a results file and choose Unroll: Open as Agent Trace. No setup, conversion or account is needed.

    TaskWithout UnrollWith Unroll
    Pick the failed runswrite a filter scriptadd a condition
    Find a tool call’s resultsearch for its IDone click
    Confirm a step matches the filefind the line by handpress R
    Open a 110 MB logless or a scriptfirst messages in under 1 s

    If one of your files doesn’t display correctly, open an issue with a small sanitized sample. That is the fastest way to get a format supported.

    — end of file · 20 records · waiting for new lines —

    Made by Mitchell Hermon. Provider and benchmark names describe compatible formats, not affiliation.

    Keyboard

    The same keys as the extension, and they work on this page too.

    J K
    Next / previous message
    N P
    Next / previous tool call
    R
    Show the raw record
    U
    Switch raw / unrolled
    Enter
    Unroll the selected raw line
    /
    Search
    F
    Follow latest
    G ⇧G
    First / last
    ⇧J ⇧K
    Next / previous run in a dataset
    ⌘S
    Save this page as unroll.jsonl
    Esc
    Clear search and filters