Each column has a name, a type and a size in bytes, written once in a header. Every row is then the same length, so a reader can calculate where any value starts instead of parsing everything before it. There is a reader in C, and a reader and writer in Python and in JavaScript.
The same three rows in both formats. JSON repeats every key in every row and puts quotes around strings. jalapenojson writes the column names once in the header and gives each column a fixed size in bytes. ␀ is the break byte, which ends a value that is shorter than its column.
The same rows again, one box per byte. All rows are the same length, so the position of a field can be calculated. Hover over a field, or move to it with Tab, to see the calculation.
This view needs JavaScript. The document above has the same bytes, with ␀ for each break byte.
しくみ
How it works
The header
The header lists each column as #name:type:size. For example, #scoville:i:7 is an integer column that is 7 bytes wide. The rows that follow contain only the values. There are no keys or quotes, nothing separates the values in a row, and nothing is escaped, so a value can contain a newline or a quote.
Fixed sizes
Every value in a column takes up exactly the column's size. A shorter value is followed by a NUL byte, the break byte, and the reader ignores the rest of the field. Sizes count bytes, not characters: jalapeño is 9 bytes because ñ is 2 bytes in UTF-8.
Finding a value
Row N, column M starts at the offset of the first row, plus N times the row length, plus the sizes of the columns before M. The reader reads that one field and nothing else. It needs no index, because it can calculate every position from the header.
A column can also hold a list of rows that use another schema. A list has no fixed size, so it is stored with its length in front of it, and a schema with a list column is read row by row instead of by offset. Lists can be nested to any depth, and the format has no limit on the number of rows or columns or on column sizes.
Column types
i integer
f float
s text
b raw bytes
y boolean
d date
t time of day
n date and time
z date and time with offset
2 list of rows using schema 2
からさ
When to use it
How it compares with JSON, based on the benchmark and on how the format works.
Good fit
Reading part of a large document, such as one row or a few columns. The reader skips the rest, while a JSON parser has to parse the whole document first. This is where the difference is largest, in every language.
Reading a whole table in C, where it is faster than yyjson, or in a JavaScript process that reads many documents, where it is faster than JSON.parse once V8 has compiled the reader.
Data stored or sent uncompressed, such as a cache, a queue or a file on disk. Column names are written once and values have no quotes or separators, so it is much smaller than the same JSON.
Nested data, such as orders with their line items. A list column keeps the size saving without flattening the data into one row per item.
Small difference
Sending it gzipped or brotli-compressed. Compression already removes most of the repeated column names, so it is only a little smaller and a little faster.
A small document loaded once by a web page. Most of the time goes to the download, and the first call to the JavaScript reader is slower than JSON.parse, so the two end up about even.
Writing, and reading every row into dicts in Python. Both take about as long as with JSON.
Poor fit
Free text. Every value in a column takes up the size of the longest one, so text that varies a lot in length wastes space.
Rows with different fields. Every row has every column, and there is no null.
Files edited by hand. A value cannot grow past its column's size without changing every row, and most editors do not show NUL bytes.
すうじ
Benchmark results
Reading and writing the same rows with JSON and with jalapenojson in each language. Every result is checked against a checksum of the rows. There are two data sets, and they are shown separately: repetitive rows, similar to typical API output, and high-entropy rows with random values that do not compress well.
Nested orders carry one to six line items each.
50,000 rows
JSON
jalapenojson
difference
repetitive, raw
5.55 MB
2.75 MB
50% smaller
repetitive, gzipped
609 KB
533 KB
13% smaller
high-entropy, raw
6.27 MB
3.50 MB
44% smaller
high-entropy, gzipped
2.27 MB
2.09 MB
8% smaller
nested orders, raw
12.04 MB
5.22 MB
57% smaller
nested orders, gzipped
1.78 MB
1.54 MB
14% smaller
Warmed up: JSON.parse and JSON.stringify against decode(), view() and encode().
50,000 rows
data
JSON
jalapenojson
difference
read every value
repetitive
24.4 ms
20.9 ms
1.2x faster
high-entropy
31.2 ms
17.2 ms
1.8x faster
read every value, as columns
repetitive
32.6 ms
17.3 ms
1.9x faster
high-entropy
39.4 ms
14.5 ms
2.7x faster
read one number column
repetitive
25.3 ms
1.69 ms
15x faster
high-entropy
31.7 ms
1.63 ms
19x faster
read one text column
repetitive
26.0 ms
4.92 ms
5.3x faster
high-entropy
33.3 ms
2.38 ms
14x faster
read one row
repetitive
23.6 ms
0.104 ms
228x faster
high-entropy
30.8 ms
0.122 ms
252x faster
write every value
repetitive
22.3 ms
15.6 ms
1.4x faster
high-entropy
22.1 ms
26.2 ms
1.2x slower
The download and the first read, which is what a page that fetches one document waits for.
50,000 rows, fast 4G (9 Mbit/s, 85 ms)
data
JSON
jalapenojson
difference
receive the document
repetitive
722 ms
577 ms
1.3x faster
high-entropy
2,139 ms
1,971 ms
1.1x faster
receive it and read every value
repetitive
751 ms
622 ms
1.2x faster
high-entropy
2,173 ms
2,012 ms
1.1x faster
receive it and read one number column
repetitive
753 ms
583 ms
1.3x faster
high-entropy
2,177 ms
1,977 ms
1.1x faster
receive it and read one row
repetitive
753 ms
580 ms
1.3x faster
high-entropy
2,173 ms
1,973 ms
1.1x faster
Warmed up: json.loads and json.dumps against the same three.
50,000 rows
data
JSON
jalapenojson
difference
read every value
repetitive
56.8 ms
50.4 ms
1.1x faster
high-entropy
54.8 ms
57.4 ms
about the same
read every value, as columns
repetitive
74.8 ms
33.5 ms
2.2x faster
high-entropy
66.0 ms
39.3 ms
1.7x faster
read one number column
repetitive
50.8 ms
7.66 ms
6.6x faster
high-entropy
44.9 ms
7.30 ms
6.1x faster
read one text column
repetitive
54.8 ms
6.47 ms
8.5x faster
high-entropy
50.0 ms
6.02 ms
8.3x faster
read one row
repetitive
59.0 ms
0.099 ms
596x faster
high-entropy
52.6 ms
0.124 ms
425x faster
write every value
repetitive
81.2 ms
57.9 ms
1.4x faster
high-entropy
86.0 ms
64.4 ms
1.3x faster
cJSON and yyjson against jj_parse() and its accessors.
50,000 rows
data
cJSON
yyjson
jalapenojson
vs cJSON
vs yyjson
read every value
repetitive
49.7 ms
6.03 ms
4.10 ms
12x faster
1.5x faster
high-entropy
54.6 ms
6.79 ms
3.87 ms
14x faster
1.8x faster
read one number column
repetitive
48.0 ms
5.72 ms
1.13 ms
43x faster
5.1x faster
high-entropy
49.9 ms
6.63 ms
1.16 ms
43x faster
5.7x faster
read one text column
repetitive
47.7 ms
5.73 ms
0.407 ms
117x faster
14x faster
high-entropy
47.9 ms
6.50 ms
0.422 ms
114x faster
15x faster
read one row
repetitive
48.1 ms
5.87 ms
0.059 ms
820x faster
100x faster
high-entropy
48.6 ms
6.54 ms
0.085 ms
574x faster
77x faster
Warmed up: the C reader compiled to WebAssembly, handing the same values to JavaScript, against JSON.parse and the JavaScript reader.
50,000 rows
data
JSON
JavaScript
WebAssembly
vs JSON
vs JavaScript
read every value
repetitive
24.4 ms
20.9 ms
20.1 ms
1.2x faster
about the same
high-entropy
31.2 ms
17.2 ms
20.8 ms
1.5x faster
1.2x slower
read every value, as columns
repetitive
32.6 ms
17.3 ms
13.6 ms
2.4x faster
1.3x faster
high-entropy
39.4 ms
14.5 ms
17.1 ms
2.3x faster
1.2x slower
read one number column
repetitive
25.3 ms
1.69 ms
1.66 ms
15x faster
about the same
high-entropy
31.7 ms
1.63 ms
1.71 ms
18x faster
1.1x slower
read one text column
repetitive
26.0 ms
4.92 ms
3.17 ms
8.2x faster
1.6x faster
high-entropy
33.3 ms
2.38 ms
3.44 ms
9.7x faster
1.4x slower
read one row
repetitive
23.6 ms
0.104 ms
0.267 ms
88x faster
2.6x slower
high-entropy
30.8 ms
0.122 ms
0.341 ms
91x faster
2.8x slower
Measured on 2026-10-03 at commit 2eb975b: Intel Xeon Processor @ 2.10GHz, 4 cores, Linux 6.18.44-fc-v64; Node 24.21.0, Python 3.11.15, cc (Ubuntu 13.3.0-6ubuntu2~24.04.1) 13.3.0 with -O2, cJSON 1.7.19, yyjson 0.13.0. The full report also has results for 1,000 rows, the first call in a new process, brotli and a slow network, and describes how each number was measured.
import { encode, decode } from "./jalapenojson.js";
const buf = encode(rows, [["id","i",4], ["name","s",16], ["seen","z",25]]);
const back = decode(buf); // Uint8Array in, objects out
back[0].seen // { instant: Date, offsetMinutes: 120 }
JavaScript
import { view } from "./jalapenojson.js";
const t = view(buf);
t.rows // how many rows, after one walk
t.get(0, "name") // one value, converted now
t.column("id") // one column; the other three are stepped over
t.row(0) // a whole row, if you want one
t.get(0, "items") // a nested list, as the bytes it is
C
jj_doc doc;
int rc = jj_parse(buf, len, &doc); // buf must outlive doc
if (rc != JJ_OK) return fprintf(stderr, "%s\n", jj_strerror(rc));
long long id = jj_int(jj_field(&doc, 0, 0));
jj_free(&doc);
The guide has more, including dates, booleans, nested lists, calculating column sizes and the rest of the C API.
Download
Each implementation is a single file with no dependencies. Copy it into your project.