jalapenojson

jalapenojson

A file format for rows of data

Each column has a name, a type and a size in bytes, written once in a header. Every row is then the same length, so a reader can calculate where any value starts instead of parsing everything before it. There is a reader in C, and a reader and writer in Python and in JavaScript.

Format version 0.6, experimental

Compared with JSON

The same three rows in both formats. JSON repeats every key in every row and puts quotes around strings. jalapenojson writes the column names once in the header and gives each column a fixed size in bytes. ␀ is the break byte, which ends a value that is shorter than its column.

JSON222 bytes minified
JSON
[
  {"pepper": "jalapeño", "scoville": 8000, "picked": "2026-09-14", "ripe": false},
  {"pepper": "serrano", "scoville": 23000, "picked": "2026-09-21", "ripe": false},
  {"pepper": "habanero", "scoville": 350000, "picked": "2026-10-02", "ripe": true}
]
jalapenojson139 bytes
jalapenojson
0.6
#pepper:s:9 #scoville:i:7 #picked:d:10 #ripe:y:1
3
jalapeño8000␀␀␀2026-09-140
serrano␀␀23000␀␀2026-09-210
habanero␀350000␀2026-10-021

How a reader finds a value

The same rows again, one box per byte. All rows are the same length, so the position of a field can be calculated. Hover over a field, or move to it with Tab, to see the calculation.

This view needs JavaScript. The document above has the same bytes, with ␀ for each break byte.

How it works

The header

The header lists each column as #name:type:size. For example, #scoville:i:7 is an integer column that is 7 bytes wide. The rows that follow contain only the values. There are no keys or quotes, nothing separates the values in a row, and nothing is escaped, so a value can contain a newline or a quote.

Fixed sizes

Every value in a column takes up exactly the column's size. A shorter value is followed by a NUL byte, the break byte, and the reader ignores the rest of the field. Sizes count bytes, not characters: jalapeño is 9 bytes because ñ is 2 bytes in UTF-8.

Finding a value

Row N, column M starts at the offset of the first row, plus N times the row length, plus the sizes of the columns before M. The reader reads that one field and nothing else. It needs no index, because it can calculate every position from the header.

A column can also hold a list of rows that use another schema. A list has no fixed size, so it is stored with its length in front of it, and a schema with a list column is read row by row instead of by offset. Lists can be nested to any depth, and the format has no limit on the number of rows or columns or on column sizes.

Column types

  • i integer
  • f float
  • s text
  • b raw bytes
  • y boolean
  • d date
  • t time of day
  • n date and time
  • z date and time with offset
  • 2 list of rows using schema 2

When to use it

How it compares with JSON, based on the benchmark and on how the format works.

Good fit

  • Reading part of a large document, such as one row or a few columns. The reader skips the rest, while a JSON parser has to parse the whole document first. This is where the difference is largest, in every language.
  • Reading a whole table in C, where it is faster than yyjson, or in a JavaScript process that reads many documents, where it is faster than JSON.parse once V8 has compiled the reader.
  • Data stored or sent uncompressed, such as a cache, a queue or a file on disk. Column names are written once and values have no quotes or separators, so it is much smaller than the same JSON.
  • Nested data, such as orders with their line items. A list column keeps the size saving without flattening the data into one row per item.

Small difference

  • Sending it gzipped or brotli-compressed. Compression already removes most of the repeated column names, so it is only a little smaller and a little faster.
  • A small document loaded once by a web page. Most of the time goes to the download, and the first call to the JavaScript reader is slower than JSON.parse, so the two end up about even.
  • Writing, and reading every row into dicts in Python. Both take about as long as with JSON.

Poor fit

  • Free text. Every value in a column takes up the size of the longest one, so text that varies a lot in length wastes space.
  • Rows with different fields. Every row has every column, and there is no null.
  • Files edited by hand. A value cannot grow past its column's size without changing every row, and most editors do not show NUL bytes.

Benchmark results

Reading and writing the same rows with JSON and with jalapenojson in each language. Every result is checked against a checksum of the rows. There are two data sets, and they are shown separately: repetitive rows, similar to typical API output, and high-entropy rows with random values that do not compress well.

Nested orders carry one to six line items each.

50,000 rowsJSONjalapenojsondifference
repetitive, raw5.55 MB2.75 MB50% smaller
repetitive, gzipped609 KB533 KB13% smaller
high-entropy, raw6.27 MB3.50 MB44% smaller
high-entropy, gzipped2.27 MB2.09 MB8% smaller
nested orders, raw12.04 MB5.22 MB57% smaller
nested orders, gzipped1.78 MB1.54 MB14% smaller

Measured on 2026-10-03 at commit 2eb975b: Intel Xeon Processor @ 2.10GHz, 4 cores, Linux 6.18.44-fc-v64; Node 24.21.0, Python 3.11.15, cc (Ubuntu 13.3.0-6ubuntu2~24.04.1) 13.3.0 with -O2, cJSON 1.7.19, yyjson 0.13.0. The full report also has results for 1,000 rows, the first call in a new process, brotli and a slow network, and describes how each number was measured.

Usage

Python
from datetime import date, datetime, timezone, timedelta
from jalapenojson import encode, decode

blob = encode(rows, schema=[("id","i",4), ("name","s",16), ("seen","z",25)])
rows = decode(blob)
rows[0]["seen"]           # datetime(2026, 1, 15, 9, 30, tzinfo=timezone(timedelta(hours=2)))

The guide has more, including dates, booleans, nested lists, calculating column sizes and the rest of the C API.

Download

Each implementation is a single file with no dependencies. Copy it into your project.