jalapenojson

jalapenojson

Every value has an address.

jalapenojson is a format for rows of data. The column names are written once, at the top, and every column takes the same number of bytes in every row. So a reader works out where any value is with one multiplication, and reads nothing else to get it.

Format 0.6 · C, Python and JavaScript · experimental

Same rows, minus the repetition

Three peppers, written twice. JSON spells out every key in every row and quotes every string. jalapenojson names the columns once, gives each one a size in bytes, and writes the values side by side. ␀ is the break byte: it ends a value that is shorter than its column.

JSON222 bytes, written compactly
JSON
[
  {"pepper": "jalapeño", "scoville": 8000, "picked": "2026-09-14", "ripe": false},
  {"pepper": "serrano", "scoville": 23000, "picked": "2026-09-21", "ripe": false},
  {"pepper": "habanero", "scoville": 350000, "picked": "2026-10-02", "ripe": true}
]
jalapenojson139 bytes
jalapenojson
0.6
#pepper:s:9 #scoville:i:7 #picked:d:10 #ripe:y:1
3
jalapeño8000␀␀␀2026-09-140
serrano␀␀23000␀␀2026-09-210
habanero␀350000␀2026-10-021

Every field, found by arithmetic

Here are the same three rows, one tile per byte. Every row is the same length, so a field's place is a sum. Point at one, or tab to it, to see the sum a reader does.

The byte-by-byte view needs JavaScript. The document above is the same bytes, with ␀ for each break byte.

How it works

Names once

The header holds each column's name, type and size: #scoville:i:7 is an integer in seven bytes. The rows hold values and nothing else. No keys, no quotes, no commas, and nothing is escaped, so a newline or a quote inside a value is just another byte.

Sizes fixed

A field takes its column's size in every row. A shorter value is followed by a break byte, a NUL, and the rest of the field is never read. A size counts bytes, not letters: jalapeño fills nine of them, because ñ takes two.

Found by arithmetic

Row N, column M starts at the first row, plus N rows, plus the sizes before M. A reader goes straight there. It builds no index, because the header already says everything an index would.

A nested list is the one column without a size: it carries its length instead, so a schema that holds one is walked rather than jumped through. Lists nest as deep as you like, and nothing else in the format has a limit either.

Nine letters and a number

  • i integer
  • f float
  • s text
  • b raw bytes
  • y yes or no
  • d date
  • t time of day
  • n date and time
  • z date and time with offset
  • 2 a list of schema 2's rows

Where it's hot, and where it's not

How much you gain over JSON, read off the benchmark and the format's own rules.

Hot

  • Reading part of a big document: one row, one column, a few columns. A reader steps over the rest in every language, while a JSON parser reads all of it first. This is where the gap is widest.
  • Reading a whole table in C, or in a JavaScript process that reads many documents.
  • Bytes stored or sent uncompressed: a cache, a queue, a file on disk. A name is written once and a value carries no quotes, commas or escapes.
  • Orders with their line items: nested lists keep the saving without flattening.

Mild

  • Over a compressed connection. gzip and brotli already squeeze out the repeated names, so it is about as much faster as it is smaller.
  • One small document fetched by a page. The download is nearly all of the wait, and a JavaScript reader's first call costs more than JSON.parse, so the two come out about the same.
  • Writing, and reading every row into dicts in Python: about what JSON costs.

Not for this

  • Free text. A column pays for its longest value in every row.
  • Ragged data. Every row has every column, and there is no null.
  • Files edited by hand. A value cannot outgrow its column without every row changing, and the break byte is a NUL that most editors will not show.

The numbers

The same rows, read and written by JSON and by jalapenojson in each language, with every result checked against a checksum of the rows. Two kinds of rows, never averaged: repetitive ones, shaped like API output, and high-entropy ones that leave a compressor little to find.

Nested orders carry one to six line items each.

50,000 rowsJSONjalapenojsondifference
repetitive, raw5.55 MB2.75 MB50% smaller
repetitive, gzipped609 KB533 KB13% smaller
high-entropy, raw6.27 MB3.50 MB44% smaller
high-entropy, gzipped2.27 MB2.09 MB8% smaller
nested orders, raw12.04 MB5.22 MB57% smaller
nested orders, gzipped1.78 MB1.54 MB14% smaller

Measured on 2026-10-03 at commit 2eb975b: Intel Xeon Processor @ 2.10GHz, 4 cores, Linux 6.18.44-fc-v64; Node 24.21.0, Python 3.11.15, cc (Ubuntu 13.3.0-6ubuntu2~24.04.1) 13.3.0 with -O2, cJSON 1.7.19, yyjson 0.13.0. The whole report also has 1,000 rows, a first call in a fresh process, brotli, a slower link, and how each number was taken.

Use it

Python
from datetime import date, datetime, timezone, timedelta
from jalapenojson import encode, decode

blob = encode(rows, schema=[("id","i",4), ("name","s",16), ("seen","z",25)])
rows = decode(blob)
rows[0]["seen"]           # datetime(2026, 1, 15, 9, 30, tzinfo=timezone(timedelta(hours=2)))

The guide covers dates, booleans, nested lists, working out column sizes, and the rest of the C API.

Get it

Each reader is one file with nothing to install. Download it and put it next to your code.