jalapenojson 0.6
A document is a list of rows. Every row has the same columns. The column names, types and sizes are written once at the top; the values follow.
Every column declares its size in bytes. A field is exactly that many bytes in every row, so a value carries no length and a reader never scans to find where a field ends. The one exception is a column holding a nested list, which carries a length instead.
Values are never escaped. Newlines, quotes and binary are all legal.
This describes format 0.6. A reader accepts 0.6 and refuses every other
version.
In the examples below,
␀stands for the break byte, which is a NUL (0x00). It is written that way so the bytes can be read on a page; a real document holds the NUL itself.python/check_spec_examples.pyputs the NULs back and reads every example here, so each one is a real document.
Layout
<major>.<minor>\n
<schema line>\n schema 1
<schema line>\n schema 2, and as many more as it needs
<row_count>\n
<rows>Then, for each row: every column of schema 1 in header order, then one \n.
A sized column is written as its field and nothing else:
<field: exactly size bytes>A list column is written as a length and then that many bytes:
<length>\n
<value: exactly length bytes>Version line
The first line. Two ASCII integers separated by .. It is always 0.6.
A reader that implements 0.6 refuses 0.5 and 0.7 alike. It does not guess.
Schema lines
One line per schema. A line is one or more column declarations separated by a single space:
#<name>:<type>:<size> a sized column
#<name>:<number> a list columnname is UTF-8. It may not contain :, a space, a newline, or NUL; any other
character is allowed, a tab included. A schema may not declare the same name
twice.
type is one of nine letters. A list column's type is instead the number of
the schema its rows use. See Types.
size is an ASCII integer of 1 or more. It is a byte count, and it has no upper
bound; see Limits.
A document declares one schema or more, as many as it needs. Schema 1 is the one the document's own rows use. The others exist to be referenced by list columns.
#id:i:4 #cur:s:3Two columns. id is 4 bytes holding an integer. cur is 3 bytes holding a
string.
Row count
An ASCII integer on the line after the last schema line. It counts the rows in this document.
It is not attached to a schema. The same schema can be used by many lists in one document, each with a different number of rows. Each nested list carries its own count.
Values
A size counts bytes, not characters.
A field is exactly its column's size in every row. A value shorter than that is followed by a break byte, which ends it; see The break byte. A value longer than the size cannot be written; an encoder refuses it. A caller with a value that does not fit decides what to do about it; the format does not decide for them.
0.6
#id:i:4 #cur:s:3
2
1001NOK
1002EURTwo rows: {id: 1001, cur: "NOK"} and {id: 1002, cur: "EUR"}. Each row is 7
bytes of field and one \n. Every value here fills its field exactly, so no
break byte appears.
The break byte
The break byte is a NUL (0x00). It marks where a value stopped inside its
field.
A value is the bytes of its field up to the first break byte, or the whole field when it holds none. What follows the break is not part of the value: a reader does not read it, check it or hand it back, so it may be anything at all. The field keeps its size either way: the break says where the value ended, not how long the field is.
0.6
#n:i:8
2
999␀␀␀␀␀
999␀1234999 in an i:8 column, twice: three bytes of value, the break, and four more
bytes. In the first row they are more break bytes; in the second they are
1234, which is not part of the value. Both rows hold 999.
So a writer never has to fill the rest of a field. The encoders in this repository fill it with break bytes anyway. That is free for a writer that builds a document in a buffer that starts out zeroed, and it keeps whatever a reused buffer held before out of the document.
No value may contain a break byte, since the first one ends it. That costs
nothing: no number, boolean, date or time grammar can hold a NUL, a s value
is UTF-8 and a NUL is not part of any text that means anything.
Padding with spaces would be simpler and is wrong, because a trailing space can
be part of a value. With a break byte, Bo and Bo in the same s:5 column
are different values and both read back as what was written.
b is the one type with no byte to reserve, since a b value may be any bytes
at all, including NULs. So a b column has no break byte: every value in it is
exactly the column's size, and an encoder refuses one that is not. A b column
holds values that are all the same length.
Types
| type | meaning | written as |
|---|---|---|
i | integer | -42 |
f | float | 3.5, 1e+300 |
s | UTF-8 string | any bytes but a NUL, up to the declared size |
b | raw bytes | any bytes, exactly the declared size |
y | boolean | 1 for true, 0 for false |
d | date | 2026-01-15 |
t | time of day | 09:30:00 or 09:30:00.123 |
n | date and time, no zone | 2026-01-15T09:30:00 |
z | date and time with offset | 2026-01-15T09:30:00+02:00 or 2026-01-15T09:30:00Z |
| a number | a list of rows of the schema with that number | see Nested lists |
Those nine letters and the numbers are every type there is. A column whose type is anything else is a malformed document: another byte, two letters, or a number with anything in it but ASCII digits. There is no unknown type that a reader reads as something else: a document written by a later version carries that version's number on its first line, and a reader refuses a version it does not know before it reaches a type, so an unknown type can only be a mistake.
Saying exactly what a type may be is also what makes readers agree about a header at all. A reader has to find where a column's name ends, where its type is and where its size begins, and readers that do that differently used to disagree about 132 of the 256 possible type bytes: a newline, a space, a second colon and every byte at or above 0x80 each found a different column depending on whether the reader scanned the line, split it, or matched it with a pattern. None of those bytes is one of the nine letters or a digit, so the question does not arise.
Numbers
i is -?[0-9]+.
f is -?[0-9]+(\.[0-9]+)?([eE][+-]?[0-9]+)?.
No leading +. No surrounding space. No underscores. No hex. No trailing
characters. Leading zeros are allowed and mean nothing, however many there are,
so 42 in an i:4 column may be written 42␀␀ or 0042 and both read back as
42.
An i value is a signed 64-bit integer, −2^63 to 2^63−1. A value outside that
range is a malformed document, not a number to round.
There is no way to write infinity or NaN. An f column always holds a real
number.
12x, 1_000, +5, 0x10 and Infinity are all errors.
Booleans
y is 1 or 0: 1 is true and 0 is false. A y value is always one
byte, so a y column declares 1.
0.6
#id:i:4 #cur:s:3 #paid:y:1
2
1001NOK1
1002EUR0Two rows: {id: 1001, cur: "NOK", paid: true} and `{id: 1002, cur: "EUR",
paid: false}`.
Nothing else is a boolean. t, true, yes, 2 and 01 are all errors, and
so is a field that holds only break bytes: that is not false, because a y
value is true or false and the format has no null.
The letter is y, for yes or no, since b already means bytes. The value is a
digit rather than a raw 0x00 or 0x01 because a NUL is the break byte, so a
false written as one would read as an empty value.
Dates
The four date types are a profile of ISO 8601, written as ASCII.
Use d when there is no time of day. Use t when there is no date. Use n for
a wall clock with no zone: it says what a clock read, not which instant. Use z
for an instant.
Fractional seconds are 1 to 6 digits. Trailing zeros are trimmed: .123, not
.123000. A zero fraction is written as nothing at all. A reader also accepts
the untrimmed spelling.
Second 60 is an error.
A d value is always 10 bytes, so a d column declares 10. The other three
vary with their fractional seconds and their zone, so their columns declare the
longest value and shorter ones are followed by a break byte.
A value that is not in the form above is an error, not a string.
docs/dates.md has the reasoning.
Row length and random access
A schema with no list column has a constant row length:
stride = sum of the column sizes + 1So a reader finds any value by arithmetic, without reading anything before it:
row N, column M = <first row's first byte> + N * stride + <sum of sizes before M>Reading one value costs the same whatever the document holds. A reader builds no index for such a schema, because there is nothing an index would record that the header does not already say. This is what the sizes are for.
A list column has no size, so a schema containing one has no constant stride and a reader walks it instead. Random access is a property of a schema, not of the format.
Nested lists
A type that is a number names a schema, counting from 1 in the order the schema lines appear. The number is ASCII digits, as many as it takes. The value in that column is a list of rows of that schema.
A list column carries a length, because a list is as long as its contents. It is the only column that does, and the only one that declares no size.
The list is written as its own row count, then its rows. It carries no header: its schema is the one the column's type names.
0.6
#order:i:4 #items:2
#sku:s:2 #qty:i:1
1
10016
1
A12
One order, 1001, carrying one item. The items value is 6 bytes: the count
1, then one row of schema 2, which is A1, 2 and the row's \n.
A schema may reference itself, which is how a tree is written. A reference may point forward, because every schema line is read before any row. Lists nest to any depth.
A reader steps over a nested list on its length like any other value. It looks inside only when asked.
More than one schema
A document declares as many as it needs, and a row may carry a list of each:
0.6
#order:i:4 #items:2 #ship:3
#sku:s:2 #qty:i:1
#city:s:6 #zip:s:4
1
100110
2
A12
B71
13
1
Oslo␀␀0150
One order carrying two items and one shipping address. Reading it down the page:
1001is the order'sid, four bytes.10is the byte length of theitemsvalue, and the ten bytes are2\n, the list's row count, thenA12\nandB71\n, two rows of schema 2.13is the byte length of theshipvalue:1\n, then one row of schema 3, which isOslo␀␀,0150and the row's\n.- The last
\nends the order's own row.
Schema 1 holds two list columns, so it has no stride and a reader walks its rows. Schemas 2 and 3 have no list column, so their rows inside each list are strided and found by arithmetic. A document can be both at once; the property belongs to each schema.
Separators
Exactly one \n after the last column of each row. A reader requires it.
A sized value carries no length and may itself begin with a \n, so the
separator is what tells a reader the row ended.
Limits
The format has none. A document holds any number of rows, a schema any number of columns, a document any number of schemas, and lists nest to any depth. A size, a row count, a length and a schema's number are as large as they need to be, and leading zeros mean nothing in any of them, however many there are.
A size larger than any document could hold is not an error. It is a claim about rows, and the one question a reader asks of it is whether the rows it describes fit in the bytes that follow. So a schema declaring one is a valid schema whose rows can never appear, and a document holding none of them reads.
The value types have ranges, which are what the types mean rather than limits
on a document: an i is a signed 64-bit integer, an f is a finite double, and
a date's year has four digits.
What bounds a document is the machine reading it: the memory it has, and the largest object its language holds at once. A reader refuses nothing else for being too large.
What a reader refuses
- A version that is not
0.6. - A column with no size that is not a list column.
- A size on a list column.
- A size that is not a positive integer.
- A column name containing
:, a space, a newline, or NUL. - Two columns with the same name in one schema.
- A type that is neither one of the nine letters nor a number.
- A number naming a schema the document does not declare. Schemas count from 1, so
0names none. - A number, boolean, date, or time that does not match its grammar.
- A row with fewer columns than the schema. Every row has every column. There are no optional fields and no null.
- A row count that does not match the bytes that follow.
- A missing separator after a row.
An encoder additionally refuses a missing value, since there is no null to write
in its place, a value longer than its column's size, a b value that is not
exactly the size, and a break byte inside a value of any type but b.
Complete example
0.6
#id:i:4 #name:s:5 #joined:d:10 #seen:z:25
2
42␀␀Alice2026-01-152026-01-15T09:30:00+02:00
7␀␀␀Bob␀␀2025-11-022025-11-02T23:05:00Z␀␀␀␀␀Two rows of 44 bytes of field and one \n. Nothing separates the values inside
a row, because each column's size says where it ends.
42 is two bytes in a four-byte field, so a break byte follows it and then one
more. Alice fills its field exactly and carries none. Bob is three bytes in
five, so it carries two. Z is five bytes shorter than what +02:00 needs, so
Bob's seen carries five.
Alice joined on a date with no time of day and was last seen at 09:30 two hours ahead of UTC. Bob was last seen at 23:05 UTC.
examples/example.jjson is this document, byte for byte, with the real NULs in
it.