jalapenojson

Guide

How to write and read jalapenojson from each language. The format itself is in the spec.

Python — date columns use the datetime types directly:

Python
from datetime import date, datetime, timezone, timedelta
from jalapenojson import encode, decode

blob = encode(rows, schema=[("id","i",4), ("name","s",16), ("seen","z",25)])
rows = decode(blob)
rows[0]["seen"]           # datetime(2026, 1, 15, 9, 30, tzinfo=timezone(timedelta(hours=2)))

d gives a date, t a time, n a naive datetime, z an aware one. On the way in, a datetime in a d column is an error rather than a silently truncated date, and a naive datetime in a z column is an error rather than a guessed offset.

A y column holds a boolean, written as 1 or 0 in one byte:

Python
blob = encode([{"id": 1001, "paid": True}], schema=[("id","i",4), ("paid","y",1)])
decode(blob)[0]["paid"]   # True

It takes True and False and nothing else, so 1, "true" and None are errors rather than guesses.

A nested list is a column whose type is a schema's number, and the schema it names is passed alongside schema 1, with as many more as you need:

Python
doc = encode(orders,
             schema=[("order","i",4), ("items","2")],
             extra_schemas=[[("sku","s",2), ("qty","i",1)]])
decode(doc)[0]["items"]   # [{"sku": "A1", "qty": 2}, ...]

The size is the third element on every column, and the encoder holds you to it — a value longer than its column is an error, not a truncated guess. A shorter one is followed by a break byte and reads back as what was written:

Python
encode(rows, schema=[("uuid","s",32), ("iso","s",3), ("cents","i",8)])

size_columns() works the sizes out from the rows you have, when you would rather not:

Python
from jalapenojson import size_columns
schema = size_columns(rows, [("uuid","s"), ("iso","s"), ("cents","i")])

JavaScript:

JavaScript
import { encode, decode } from "./jalapenojson.js";
const buf = encode(rows, [["id","i",4], ["name","s",16], ["seen","z",25]]);
const back = decode(buf);          // Uint8Array in, objects out
back[0].seen                       // { instant: Date, offsetMinutes: 120 }

d, t and n come back as canonical strings, because JavaScript has no type for a date without a zone. z comes back as { instant, offsetMinutes, micros }: a JS Date is an instant, so it carries neither the offset — which would leave the value re-rendered in the reader's timezone — nor a fraction finer than a millisecond. encode accepts that shape, a canonical string, or a bare Date (written against UTC, to millisecond precision). A y column comes back as true or false, and encode takes those two and nothing else.

Nested lists take the schema they name as a third argument, and come back as arrays of objects:

JavaScript
const buf = encode(orders, [["order","i",4], ["items","2"]],
                   [[["sku","s",2], ["qty","i",1]]]);
decode(buf)[0].items               // [{ sku: "A1", qty: 2 }, ...]

The size is the third element on every column, the same as in Python, and sizeColumns() works them out from your rows:

JavaScript
encode(rows, [["uuid","s",32], ["iso","s",3], ["cents","i",8]]);
sizeColumns(rows, [["uuid","s"], ["iso","s"], ["cents","i"]]);

view() reads the same document without building it. A row is a constant number of bytes, so a value's address is arithmetic and nothing at all is walked or recorded; it converts only the values it is asked for:

JavaScript
import { view } from "./jalapenojson.js";
const t = view(buf);
t.rows                             // how many rows, after one walk
t.get(0, "name")                   // one value, converted now
t.column("id")                     // one column; the other three are stepped over
t.row(0)                           // a whole row, if you want one
t.get(0, "items")                  // a nested list, as the bytes it is

A nested column is the one view() does not convert: building the list would be the materializing that view() exists to avoid, so get() hands back the value and decode() on it is the caller's own step. This is not a faster way to do the same work — it is a way to do less of it. Use view() when you want part of a document, or its columns as arrays, and decode() when you want every row as an object.

C: c/jalapenojson.h is a single-header zero-copy parser.

C
jj_doc doc;
int rc = jj_parse(buf, len, &doc);          // buf must outlive doc
if (rc != JJ_OK) return fprintf(stderr, "%s\n", jj_strerror(rc));
long long id = jj_int(jj_field(&doc, 0, 0));
jj_free(&doc);

jj_parse mutates buf (column names are NUL-terminated in place) and validates every length against len, so the buffer does not need to be NUL-terminated. It does not read a value's bytes, so a value comes back as the bytes it is. For a document you did not write, jj_int_checked and jj_float_checked say whether a number is one, jj_bool_checked whether a y value is 1 or 0, and jj_str_checked whether an s value is UTF-8. jj_parse does not descend into a nested list: it records the value as the bytes it is, which costs nothing, and jj_sub reads one when you ask.

C
if (jj_is_list(doc.col_types[1])) {
    jj_doc items;
    // col_schemas says which schema the list's rows use.
    if (jj_sub(&doc, doc.col_schemas[1], jj_field(&doc, 0, 1), &items) == JJ_OK) {
        // items is an ordinary jj_doc over the same buffer, and its own
        // columns may hold further lists, read the same way.
        jj_free(&items);
    }
}

Because the parser never recurses, a document cannot drive it off the stack; how deep to go is the caller's own decision. Python and JavaScript read every list and do not recurse either, so lists nest as deep as memory allows in all three.

doc.col_sizes[c] is a column's declared size, doc.col_offs[c] is where it starts inside a row, and doc.stride is how long a row is. When doc.stride is non-zero, doc.vals is NULL: no index was built, because jj_field works a value's address out from those three. A schema holding a nested list has no constant row length, so there doc.stride is 0 and the rows were walked into doc.vals instead. Either way jj_field means the same thing, and a value's len stops at its break byte, so it means the value in every implementation.

Dates live in c/jj_date.h, layered on top and entirely optional:

C
#include "jj_date.h"
jj_date d;
if (jj_date_parse(jj_field(&doc, r, c), doc.col_types[c], &d) == JJ_DATE_OK)
    printf("%lld\n", jj_date_epoch_sec(&d));   // needs a 'z' column