Guide
How to write and read jalapenojson from each language. The format itself is in the spec.
Python — date columns use the datetime types directly:
from datetime import date, datetime, timezone, timedelta
from jalapenojson import encode, decode
blob = encode(rows, schema=[("id","i",4), ("name","s",16), ("seen","z",25)])
rows = decode(blob)
rows[0]["seen"] # datetime(2026, 1, 15, 9, 30, tzinfo=timezone(timedelta(hours=2)))d gives a date, t a time, n a naive datetime, z an aware one. On
the way in, a datetime in a d column is an error rather than a silently
truncated date, and a naive datetime in a z column is an error rather than a
guessed offset.
A y column holds a boolean, written as 1 or 0 in one byte:
blob = encode([{"id": 1001, "paid": True}], schema=[("id","i",4), ("paid","y",1)])
decode(blob)[0]["paid"] # TrueIt takes True and False and nothing else, so 1, "true" and None are
errors rather than guesses.
A nested list is a column whose type is a schema's number, and the schema it names is passed alongside schema 1, with as many more as you need:
doc = encode(orders,
schema=[("order","i",4), ("items","2")],
extra_schemas=[[("sku","s",2), ("qty","i",1)]])
decode(doc)[0]["items"] # [{"sku": "A1", "qty": 2}, ...]The size is the third element on every column, and the encoder holds you to it — a value longer than its column is an error, not a truncated guess. A shorter one is followed by a break byte and reads back as what was written:
encode(rows, schema=[("uuid","s",32), ("iso","s",3), ("cents","i",8)])size_columns() works the sizes out from the rows you have, when you would
rather not:
from jalapenojson import size_columns
schema = size_columns(rows, [("uuid","s"), ("iso","s"), ("cents","i")])JavaScript:
import { encode, decode } from "./jalapenojson.js";
const buf = encode(rows, [["id","i",4], ["name","s",16], ["seen","z",25]]);
const back = decode(buf); // Uint8Array in, objects out
back[0].seen // { instant: Date, offsetMinutes: 120 }d, t and n come back as canonical strings, because JavaScript has no type
for a date without a zone. z comes back as { instant, offsetMinutes, micros }:
a JS Date is an instant, so it carries neither the offset — which would leave
the value re-rendered in the reader's timezone — nor a fraction finer than a
millisecond. encode accepts that shape, a canonical string, or a bare Date
(written against UTC, to millisecond precision). A y column comes back as
true or false, and encode takes those two and nothing else.
Nested lists take the schema they name as a third argument, and come back as arrays of objects:
const buf = encode(orders, [["order","i",4], ["items","2"]],
[[["sku","s",2], ["qty","i",1]]]);
decode(buf)[0].items // [{ sku: "A1", qty: 2 }, ...]The size is the third element on every column, the same as in Python, and
sizeColumns() works them out from your rows:
encode(rows, [["uuid","s",32], ["iso","s",3], ["cents","i",8]]);
sizeColumns(rows, [["uuid","s"], ["iso","s"], ["cents","i"]]);view() reads the same document without building it. A row is a constant
number of bytes, so a value's address is arithmetic and nothing at all is
walked or recorded; it converts only the values it is asked for:
import { view } from "./jalapenojson.js";
const t = view(buf);
t.rows // how many rows, after one walk
t.get(0, "name") // one value, converted now
t.column("id") // one column; the other three are stepped over
t.row(0) // a whole row, if you want one
t.get(0, "items") // a nested list, as the bytes it isA nested column is the one view() does not convert: building the list would
be the materializing that view() exists to avoid, so get() hands back the
value and decode() on it is the caller's own step.
This is not a faster way to do the same work — it is a way to do less of it.
Use view() when you want part of a document, or its columns as arrays, and
decode() when you want every row as an object.
C: c/jalapenojson.h is a single-header zero-copy parser.
jj_doc doc;
int rc = jj_parse(buf, len, &doc); // buf must outlive doc
if (rc != JJ_OK) return fprintf(stderr, "%s\n", jj_strerror(rc));
long long id = jj_int(jj_field(&doc, 0, 0));
jj_free(&doc);jj_parse mutates buf (column names are NUL-terminated in place) and validates
every length against len, so the buffer does not need to be NUL-terminated.
It does not read a value's bytes, so a value comes back as the bytes it is. For
a document you did not write, jj_int_checked and jj_float_checked say
whether a number is one, jj_bool_checked whether a y value is 1 or 0,
and jj_str_checked whether an s value is UTF-8.
jj_parse does not descend into a nested list: it records the value as the
bytes it is, which costs nothing, and jj_sub reads one when you ask.
if (jj_is_list(doc.col_types[1])) {
jj_doc items;
// col_schemas says which schema the list's rows use.
if (jj_sub(&doc, doc.col_schemas[1], jj_field(&doc, 0, 1), &items) == JJ_OK) {
// items is an ordinary jj_doc over the same buffer, and its own
// columns may hold further lists, read the same way.
jj_free(&items);
}
}Because the parser never recurses, a document cannot drive it off the stack; how deep to go is the caller's own decision. Python and JavaScript read every list and do not recurse either, so lists nest as deep as memory allows in all three.
doc.col_sizes[c] is a column's declared size, doc.col_offs[c] is where it
starts inside a row, and doc.stride is how long a row is. When doc.stride is
non-zero, doc.vals is NULL: no index was built, because jj_field works a
value's address out from those three. A schema holding a nested list has no
constant row length, so there doc.stride is 0 and the rows were walked into
doc.vals instead. Either way jj_field means the same thing, and a value's
len stops at its break byte, so it means the value in every implementation.
Dates live in c/jj_date.h, layered on top and entirely optional:
#include "jj_date.h"
jj_date d;
if (jj_date_parse(jj_field(&doc, r, c), doc.col_types[c], &d) == JJ_DATE_OK)
printf("%lld\n", jj_date_epoch_sec(&d)); // needs a 'z' column