jalapenojson is a format for rows of data. The column names are written once, at the top, and every column takes the same number of bytes in every row. So a reader works out where any value is with one multiplication, and reads nothing else to get it.
Format 0.6 · C, Python and JavaScript · experimental
くらべる
Same rows, minus the repetition
Three peppers, written twice. JSON spells out every key in every row and quotes every string. jalapenojson names the columns once, gives each one a size in bytes, and writes the values side by side. ␀ is the break byte: it ends a value that is shorter than its column.
Here are the same three rows, one tile per byte. Every row is the same length, so a field's place is a sum. Point at one, or tab to it, to see the sum a reader does.
The byte-by-byte view needs JavaScript. The document above is the same bytes, with ␀ for each break byte.
しくみ
How it works
Names once
The header holds each column's name, type and size: #scoville:i:7 is an integer in seven bytes. The rows hold values and nothing else. No keys, no quotes, no commas, and nothing is escaped, so a newline or a quote inside a value is just another byte.
Sizes fixed
A field takes its column's size in every row. A shorter value is followed by a break byte, a NUL, and the rest of the field is never read. A size counts bytes, not letters: jalapeño fills nine of them, because ñ takes two.
Found by arithmetic
Row N, column M starts at the first row, plus N rows, plus the sizes before M. A reader goes straight there. It builds no index, because the header already says everything an index would.
A nested list is the one column without a size: it carries its length instead, so a schema that holds one is walked rather than jumped through. Lists nest as deep as you like, and nothing else in the format has a limit either.
Nine letters and a number
i integer
f float
s text
b raw bytes
y yes or no
d date
t time of day
n date and time
z date and time with offset
2 a list of schema 2's rows
からさ
Where it's hot, and where it's not
How much you gain over JSON, read off the benchmark and the format's own rules.
Hot
Reading part of a big document: one row, one column, a few columns. A reader steps over the rest in every language, while a JSON parser reads all of it first. This is where the gap is widest.
Reading a whole table in C, or in a JavaScript process that reads many documents.
Bytes stored or sent uncompressed: a cache, a queue, a file on disk. A name is written once and a value carries no quotes, commas or escapes.
Orders with their line items: nested lists keep the saving without flattening.
Mild
Over a compressed connection. gzip and brotli already squeeze out the repeated names, so it is about as much faster as it is smaller.
One small document fetched by a page. The download is nearly all of the wait, and a JavaScript reader's first call costs more than JSON.parse, so the two come out about the same.
Writing, and reading every row into dicts in Python: about what JSON costs.
Not for this
Free text. A column pays for its longest value in every row.
Ragged data. Every row has every column, and there is no null.
Files edited by hand. A value cannot outgrow its column without every row changing, and the break byte is a NUL that most editors will not show.
すうじ
The numbers
The same rows, read and written by JSON and by jalapenojson in each language, with every result checked against a checksum of the rows. Two kinds of rows, never averaged: repetitive ones, shaped like API output, and high-entropy ones that leave a compressor little to find.
Nested orders carry one to six line items each.
50,000 rows
JSON
jalapenojson
difference
repetitive, raw
5.55 MB
2.75 MB
50% smaller
repetitive, gzipped
609 KB
533 KB
13% smaller
high-entropy, raw
6.27 MB
3.50 MB
44% smaller
high-entropy, gzipped
2.27 MB
2.09 MB
8% smaller
nested orders, raw
12.04 MB
5.22 MB
57% smaller
nested orders, gzipped
1.78 MB
1.54 MB
14% smaller
Warmed up: JSON.parse and JSON.stringify against decode(), view() and encode().
50,000 rows
data
JSON
jalapenojson
difference
read every value
repetitive
24.4 ms
20.9 ms
1.2x faster
high-entropy
31.2 ms
17.2 ms
1.8x faster
read every value, as columns
repetitive
32.6 ms
17.3 ms
1.9x faster
high-entropy
39.4 ms
14.5 ms
2.7x faster
read one number column
repetitive
25.3 ms
1.69 ms
15x faster
high-entropy
31.7 ms
1.63 ms
19x faster
read one text column
repetitive
26.0 ms
4.92 ms
5.3x faster
high-entropy
33.3 ms
2.38 ms
14x faster
read one row
repetitive
23.6 ms
0.104 ms
228x faster
high-entropy
30.8 ms
0.122 ms
252x faster
write every value
repetitive
22.3 ms
15.6 ms
1.4x faster
high-entropy
22.1 ms
26.2 ms
1.2x slower
The download and the first read, which is what a page that fetches one document waits for.
50,000 rows, fast 4G (9 Mbit/s, 85 ms)
data
JSON
jalapenojson
difference
receive the document
repetitive
722 ms
577 ms
1.3x faster
high-entropy
2,139 ms
1,971 ms
1.1x faster
receive it and read every value
repetitive
751 ms
622 ms
1.2x faster
high-entropy
2,173 ms
2,012 ms
1.1x faster
receive it and read one number column
repetitive
753 ms
583 ms
1.3x faster
high-entropy
2,177 ms
1,977 ms
1.1x faster
receive it and read one row
repetitive
753 ms
580 ms
1.3x faster
high-entropy
2,173 ms
1,973 ms
1.1x faster
Warmed up: json.loads and json.dumps against the same three.
50,000 rows
data
JSON
jalapenojson
difference
read every value
repetitive
56.8 ms
50.4 ms
1.1x faster
high-entropy
54.8 ms
57.4 ms
about the same
read every value, as columns
repetitive
74.8 ms
33.5 ms
2.2x faster
high-entropy
66.0 ms
39.3 ms
1.7x faster
read one number column
repetitive
50.8 ms
7.66 ms
6.6x faster
high-entropy
44.9 ms
7.30 ms
6.1x faster
read one text column
repetitive
54.8 ms
6.47 ms
8.5x faster
high-entropy
50.0 ms
6.02 ms
8.3x faster
read one row
repetitive
59.0 ms
0.099 ms
596x faster
high-entropy
52.6 ms
0.124 ms
425x faster
write every value
repetitive
81.2 ms
57.9 ms
1.4x faster
high-entropy
86.0 ms
64.4 ms
1.3x faster
cJSON and yyjson against jj_parse() and its accessors.
50,000 rows
data
cJSON
yyjson
jalapenojson
vs cJSON
vs yyjson
read every value
repetitive
49.7 ms
6.03 ms
4.10 ms
12x faster
1.5x faster
high-entropy
54.6 ms
6.79 ms
3.87 ms
14x faster
1.8x faster
read one number column
repetitive
48.0 ms
5.72 ms
1.13 ms
43x faster
5.1x faster
high-entropy
49.9 ms
6.63 ms
1.16 ms
43x faster
5.7x faster
read one text column
repetitive
47.7 ms
5.73 ms
0.407 ms
117x faster
14x faster
high-entropy
47.9 ms
6.50 ms
0.422 ms
114x faster
15x faster
read one row
repetitive
48.1 ms
5.87 ms
0.059 ms
820x faster
100x faster
high-entropy
48.6 ms
6.54 ms
0.085 ms
574x faster
77x faster
Warmed up: the C reader compiled to WebAssembly, handing the same values to JavaScript, against JSON.parse and the JavaScript reader.
50,000 rows
data
JSON
JavaScript
WebAssembly
vs JSON
vs JavaScript
read every value
repetitive
24.4 ms
20.9 ms
20.1 ms
1.2x faster
about the same
high-entropy
31.2 ms
17.2 ms
20.8 ms
1.5x faster
1.2x slower
read every value, as columns
repetitive
32.6 ms
17.3 ms
13.6 ms
2.4x faster
1.3x faster
high-entropy
39.4 ms
14.5 ms
17.1 ms
2.3x faster
1.2x slower
read one number column
repetitive
25.3 ms
1.69 ms
1.66 ms
15x faster
about the same
high-entropy
31.7 ms
1.63 ms
1.71 ms
18x faster
1.1x slower
read one text column
repetitive
26.0 ms
4.92 ms
3.17 ms
8.2x faster
1.6x faster
high-entropy
33.3 ms
2.38 ms
3.44 ms
9.7x faster
1.4x slower
read one row
repetitive
23.6 ms
0.104 ms
0.267 ms
88x faster
2.6x slower
high-entropy
30.8 ms
0.122 ms
0.341 ms
91x faster
2.8x slower
Measured on 2026-10-03 at commit 2eb975b: Intel Xeon Processor @ 2.10GHz, 4 cores, Linux 6.18.44-fc-v64; Node 24.21.0, Python 3.11.15, cc (Ubuntu 13.3.0-6ubuntu2~24.04.1) 13.3.0 with -O2, cJSON 1.7.19, yyjson 0.13.0. The whole report also has 1,000 rows, a first call in a fresh process, brotli, a slower link, and how each number was taken.
import { encode, decode } from "./jalapenojson.js";
const buf = encode(rows, [["id","i",4], ["name","s",16], ["seen","z",25]]);
const back = decode(buf); // Uint8Array in, objects out
back[0].seen // { instant: Date, offsetMinutes: 120 }
JavaScript
import { view } from "./jalapenojson.js";
const t = view(buf);
t.rows // how many rows, after one walk
t.get(0, "name") // one value, converted now
t.column("id") // one column; the other three are stepped over
t.row(0) // a whole row, if you want one
t.get(0, "items") // a nested list, as the bytes it is
C
jj_doc doc;
int rc = jj_parse(buf, len, &doc); // buf must outlive doc
if (rc != JJ_OK) return fprintf(stderr, "%s\n", jj_strerror(rc));
long long id = jj_int(jj_field(&doc, 0, 0));
jj_free(&doc);
The guide covers dates, booleans, nested lists, working out column sizes, and the rest of the C API.
Get it
Each reader is one file with nothing to install. Download it and put it next to your code.