Anatomy of one formula Every read that one =ES.Get() causes — byte offsets, lengths, milliseconds — measured, not guessed.

Sparrow's folder mode points Excel at a directory of data files — no server, no database, no session. Fair question: what actually happens when you press Enter? We instrumented the add-in's own readers and traced one formula asking for one monthly series — 125 rows out of a 14 MB, 3.2-million-row file sitting on an ordinary Windows network share. Everything below is a measurement from that trace.

The whole journey: ~234 ms, and where they go

One formula, warm file. The long amber block is deliberate: a 200 ms batching window in which the add-in waits for sibling formulas. A lone formula pays it in full — a hundred formulas in a refresh pay it once, together, then share a single file open. That trade is the add-in's core design: the batch of 100 lands in ~280 ms total, about 3 ms per formula.

After the window fires, the engine's entire contribution is ~28 ms — and most of that is CPU (decoding compressed columns), not I/O. The interesting part is how little I/O there is, and where it lands.

Where the reads land

Both charts below plot the same query on the same clock: x = time, y = byte offset in the file (offset 0 at the top, end of file at the bottom). Every mark is one actual read call, captured live.

re-draws both charts in time order, 76× slow motion — same data, same code

Vortex — 4 reads, done in 8 ms

One 64 KB read of the file's tail (teal, bottom) fetches the footer: schema, per-chunk statistics, and a complete map of every segment's byte range — 5.7 KB of metadata that describes all 14 MB. The engine prunes chunks whose statistics exclude the series, then issues the three surviving ranges in parallel — they all start at the same instant, because nothing about them was left to discover. Then the chart goes quiet: there is nothing left to read.

Parquet — 5,914 reads, 104 ms

The same 125 rows through the add-in's parquet path. Each dot is one read (dot area ∝ bytes). The footer tells the reader which row groups to skip and where each column chunk begins — but inside a chunk, data is a chain of pages, and each page's header must be read to learn where the next page starts. The dots form one unbroken causal thread: read n contains the address of read n+1, so no read can begin until the previous one returns.

The format invites chatty reading — and most readers accept the invitation.

The same two traces, raced

Both charts above on one clock, 85× slow motion. The teal panel is finished — 4 reads, 8 ms — before the parquet chain clears its first pages; the counter keeps climbing for another 96 ms. Every dot is one captured read.

To be fair to parquet, we ran the same query through DuckDB — one of the best parquet readers there is. It pruned tighter (1.26 MB touched) but still issued ~192 reads. The chattiness isn't a library quirk; it's the shape of the format. Parquet was designed in the Hadoop era for distributed full scans — read whole column chunks of whole row groups, throughput over latency. It tolerates database-style queries ("give me this one series"). Vortex is the same niche redesigned a decade later, after the world learned that queries against files look like database queries: its footer is not a table of contents but a complete byte map, and its encodings decode selectively. Same admission price — one footer read — entirely different amount of knowledge purchased.

Parquet still earns its place: it's readable by every tool on earth, and on our data it's consistently about half the size on disk. That's why the add-in reads both and simply prefers .vortex when a dataset ships it.

What batching buys

The remarkable property of the planned-read design: the read count doesn't grow with the batch. We ran a real analyst watch list at four sizes against an 18.5 MB file over the same share:

batchrows returnedfile readsbytes touchedwarm query
1 series2,68132.0 MB29 ms
10 series5,55023.6 MB26 ms
100 series25,81049.2 MB45 ms
all 1,391215,231317.0 MB~510 ms

One series or the entire watch list: always a footer plus two or three coalesced ranges. Batching fattens the reads; it never multiplies them. Combined with the 200 ms batching window, a whole-workbook refresh — 1,391 series, 215 thousand rows — is one file open, three reads, and about half a second. From a plain network share, with no server involved.

Why this matters for a workbook

Every layer underneath cooperates with this pattern. Windows caches what was read, so repeat refreshes move zero bytes over the wire until the file actually changes (we measured it: 0.00 MB). SMB grants read leases, so a hundred analysts pointed at the same shared folder each pull the ranges once per file version and then serve themselves from local cache — the file server does almost nothing. And when the publisher replaces a file, the replacement is atomic: every reader sees the old version or the new one, never a torn middle.

That's the folder-mode bet in one page: if reads are planned, fat, and parallel, a directory of well-built files behaves like a very fast, very boring database — one with no server to run, no sessions to manage, and nothing to go down.

Coda: the same plan, over the internet

Because the reads are planned, the transport underneath barely matters. We pointed the same engine at the same file served over HTTPS from object storage (Cloudflare R2) — no local copy, no download. The entire query was four requests:

HEAD                                   (file size)
GET  bytes 18,392,309 – 18,457,843     (64 KB — the footer)
GET  bytes 14,757,824 – 18,419,939     (3.6 MB)
GET  bytes 11,844,448 – 12,396,843     (539 KB)

Those two data ranges are byte-identical to the local reads in the vortex chart above — same footer, same map, same coalesced ranges, whether they become file reads or HTTP range requests. One series over a real mobile-broadband WAN: ~0.6–1.2 s cold. Repeat queries on the open file answered from the engine's in-memory segment cache in 1–2 ms. The full 100-series batch: 2.1 s.

And it isn't our reader that makes this work. Vortex is a DuckDB core extension, and we pointed a stock DuckDB — an engine we didn't write, on a machine with none of our software — at the same published URL: it answered SQL against the same file. The format is open, the folder specification is open, and any engine that speaks either can serve itself.

Today the add-in handles cloud sources by downloading the file once and revalidating it cheaply — the right trade for a workbook refreshed all day. These measurements are the road ahead: a plan built from one footer read works identically against a disk, a network share, or an object store on the other side of the world.

Try it — the EIA Data Edition is free   Read the folder specification

How we measured it

Every number on this page comes from the add-in's shipped readers, not a synthetic benchmark: a console harness compiles the same vortex_ffi binding and the same Parquet.Net path the product runs, and traces one formula's work. Four instrumentation layers cross-check each other:

For the HTTP legs, a range-logging proxy plays the server's role. Every "cold" run read a freshly written server-side copy, so neither the client cache nor the server cache could quietly warm it; "warm" means the OS cache holds the file, and cold adds one wire transfer of the touched bytes. DuckDB 1.5.4 ran the same query as the independent control. The charts above are drawn from the captured (time, offset, length) tuples directly — the raw read log ships inside this page's source.

Measured 2026-08-23: public EIA petroleum data over an SMB share with ~1 ms round-trip.