jalapenojson

jalapenojson 0.6

A document is a list of rows. Every row has the same columns. The column names, types and sizes are written once at the top; the values follow.

Every column declares its size in bytes. A field is exactly that many bytes in every row, so a value carries no length and a reader never scans to find where a field ends. The one exception is a column holding a nested list, which carries a length instead.

Values are never escaped. Newlines, quotes and binary are all legal.

This describes format 0.6. A reader accepts 0.6 and refuses every other version.

In the examples below, ␀ stands for the break byte, which is a NUL (0x00). It is written that way so the bytes can be read on a page; a real document holds the NUL itself. python/check_spec_examples.py puts the NULs back and reads every example here, so each one is a real document.

Layout

<major>.<minor>\n
<schema line>\n          schema 1
<schema line>\n          schema 2, and as many more as it needs
<row_count>\n
<rows>

Then, for each row: every column of schema 1 in header order, then one \n.

A sized column is written as its field and nothing else:

<field: exactly size bytes>

A list column is written as a length and then that many bytes:

<length>\n
<value: exactly length bytes>

Version line

The first line. Two ASCII integers separated by .. It is always 0.6.

A reader that implements 0.6 refuses 0.5 and 0.7 alike. It does not guess.

Schema lines

One line per schema. A line is one or more column declarations separated by a single space:

#<name>:<type>:<size>     a sized column
#<name>:<number>          a list column

name is UTF-8. It may not contain :, a space, a newline, or NUL; any other character is allowed, a tab included. A schema may not declare the same name twice.

type is one of nine letters. A list column's type is instead the number of the schema its rows use. See Types.

size is an ASCII integer of 1 or more. It is a byte count, and it has no upper bound; see Limits.

A document declares one schema or more, as many as it needs. Schema 1 is the one the document's own rows use. The others exist to be referenced by list columns.

#id:i:4 #cur:s:3

Two columns. id is 4 bytes holding an integer. cur is 3 bytes holding a string.

Row count

An ASCII integer on the line after the last schema line. It counts the rows in this document.

It is not attached to a schema. The same schema can be used by many lists in one document, each with a different number of rows. Each nested list carries its own count.

Values

A size counts bytes, not characters.

A field is exactly its column's size in every row. A value shorter than that is followed by a break byte, which ends it; see The break byte. A value longer than the size cannot be written; an encoder refuses it. A caller with a value that does not fit decides what to do about it; the format does not decide for them.

jalapenojson
0.6
#id:i:4 #cur:s:3
2
1001NOK
1002EUR

Two rows: {id: 1001, cur: "NOK"} and {id: 1002, cur: "EUR"}. Each row is 7 bytes of field and one \n. Every value here fills its field exactly, so no break byte appears.

The break byte

The break byte is a NUL (0x00). It marks where a value stopped inside its field.

A value is the bytes of its field up to the first break byte, or the whole field when it holds none. What follows the break is not part of the value: a reader does not read it, check it or hand it back, so it may be anything at all. The field keeps its size either way: the break says where the value ended, not how long the field is.

jalapenojson
0.6
#n:i:8
2
999␀␀␀␀␀
999␀1234

999 in an i:8 column, twice: three bytes of value, the break, and four more bytes. In the first row they are more break bytes; in the second they are 1234, which is not part of the value. Both rows hold 999.

So a writer never has to fill the rest of a field. The encoders in this repository fill it with break bytes anyway. That is free for a writer that builds a document in a buffer that starts out zeroed, and it keeps whatever a reused buffer held before out of the document.

No value may contain a break byte, since the first one ends it. That costs nothing: no number, boolean, date or time grammar can hold a NUL, a s value is UTF-8 and a NUL is not part of any text that means anything.

Padding with spaces would be simpler and is wrong, because a trailing space can be part of a value. With a break byte, Bo and Bo in the same s:5 column are different values and both read back as what was written.

b is the one type with no byte to reserve, since a b value may be any bytes at all, including NULs. So a b column has no break byte: every value in it is exactly the column's size, and an encoder refuses one that is not. A b column holds values that are all the same length.

Types

typemeaningwritten as
iinteger-42
ffloat3.5, 1e+300
sUTF-8 stringany bytes but a NUL, up to the declared size
braw bytesany bytes, exactly the declared size
yboolean1 for true, 0 for false
ddate2026-01-15
ttime of day09:30:00 or 09:30:00.123
ndate and time, no zone2026-01-15T09:30:00
zdate and time with offset2026-01-15T09:30:00+02:00 or 2026-01-15T09:30:00Z
a numbera list of rows of the schema with that numbersee Nested lists

Those nine letters and the numbers are every type there is. A column whose type is anything else is a malformed document: another byte, two letters, or a number with anything in it but ASCII digits. There is no unknown type that a reader reads as something else: a document written by a later version carries that version's number on its first line, and a reader refuses a version it does not know before it reaches a type, so an unknown type can only be a mistake.

Saying exactly what a type may be is also what makes readers agree about a header at all. A reader has to find where a column's name ends, where its type is and where its size begins, and readers that do that differently used to disagree about 132 of the 256 possible type bytes: a newline, a space, a second colon and every byte at or above 0x80 each found a different column depending on whether the reader scanned the line, split it, or matched it with a pattern. None of those bytes is one of the nine letters or a digit, so the question does not arise.

Numbers

i is -?[0-9]+.

f is -?[0-9]+(\.[0-9]+)?([eE][+-]?[0-9]+)?.

No leading +. No surrounding space. No underscores. No hex. No trailing characters. Leading zeros are allowed and mean nothing, however many there are, so 42 in an i:4 column may be written 42␀␀ or 0042 and both read back as 42.

An i value is a signed 64-bit integer, −2^63 to 2^63−1. A value outside that range is a malformed document, not a number to round.

There is no way to write infinity or NaN. An f column always holds a real number.

12x, 1_000, +5, 0x10 and Infinity are all errors.

Booleans

y is 1 or 0: 1 is true and 0 is false. A y value is always one byte, so a y column declares 1.

jalapenojson
0.6
#id:i:4 #cur:s:3 #paid:y:1
2
1001NOK1
1002EUR0

Two rows: {id: 1001, cur: "NOK", paid: true} and `{id: 1002, cur: "EUR", paid: false}`.

Nothing else is a boolean. t, true, yes, 2 and 01 are all errors, and so is a field that holds only break bytes: that is not false, because a y value is true or false and the format has no null.

The letter is y, for yes or no, since b already means bytes. The value is a digit rather than a raw 0x00 or 0x01 because a NUL is the break byte, so a false written as one would read as an empty value.

Dates

The four date types are a profile of ISO 8601, written as ASCII.

Use d when there is no time of day. Use t when there is no date. Use n for a wall clock with no zone: it says what a clock read, not which instant. Use z for an instant.

Fractional seconds are 1 to 6 digits. Trailing zeros are trimmed: .123, not .123000. A zero fraction is written as nothing at all. A reader also accepts the untrimmed spelling.

Second 60 is an error.

A d value is always 10 bytes, so a d column declares 10. The other three vary with their fractional seconds and their zone, so their columns declare the longest value and shorter ones are followed by a break byte.

A value that is not in the form above is an error, not a string.

docs/dates.md has the reasoning.

Row length and random access

A schema with no list column has a constant row length:

stride = sum of the column sizes + 1

So a reader finds any value by arithmetic, without reading anything before it:

row N, column M  =  <first row's first byte> + N * stride + <sum of sizes before M>

Reading one value costs the same whatever the document holds. A reader builds no index for such a schema, because there is nothing an index would record that the header does not already say. This is what the sizes are for.

A list column has no size, so a schema containing one has no constant stride and a reader walks it instead. Random access is a property of a schema, not of the format.

Nested lists

A type that is a number names a schema, counting from 1 in the order the schema lines appear. The number is ASCII digits, as many as it takes. The value in that column is a list of rows of that schema.

A list column carries a length, because a list is as long as its contents. It is the only column that does, and the only one that declares no size.

The list is written as its own row count, then its rows. It carries no header: its schema is the one the column's type names.

jalapenojson
0.6
#order:i:4 #items:2
#sku:s:2 #qty:i:1
1
10016
1
A12

One order, 1001, carrying one item. The items value is 6 bytes: the count 1, then one row of schema 2, which is A1, 2 and the row's \n.

A schema may reference itself, which is how a tree is written. A reference may point forward, because every schema line is read before any row. Lists nest to any depth.

A reader steps over a nested list on its length like any other value. It looks inside only when asked.

More than one schema

A document declares as many as it needs, and a row may carry a list of each:

jalapenojson
0.6
#order:i:4 #items:2 #ship:3
#sku:s:2 #qty:i:1
#city:s:6 #zip:s:4
1
100110
2
A12
B71
13
1
Oslo␀␀0150

One order carrying two items and one shipping address. Reading it down the page:

  • 1001 is the order's id, four bytes.
  • 10 is the byte length of the items value, and the ten bytes are 2\n, the list's row count, then A12\n and B71\n, two rows of schema 2.
  • 13 is the byte length of the ship value: 1\n, then one row of schema 3, which is Oslo␀␀, 0150 and the row's \n.
  • The last \n ends the order's own row.

Schema 1 holds two list columns, so it has no stride and a reader walks its rows. Schemas 2 and 3 have no list column, so their rows inside each list are strided and found by arithmetic. A document can be both at once; the property belongs to each schema.

Separators

Exactly one \n after the last column of each row. A reader requires it.

A sized value carries no length and may itself begin with a \n, so the separator is what tells a reader the row ended.

Limits

The format has none. A document holds any number of rows, a schema any number of columns, a document any number of schemas, and lists nest to any depth. A size, a row count, a length and a schema's number are as large as they need to be, and leading zeros mean nothing in any of them, however many there are.

A size larger than any document could hold is not an error. It is a claim about rows, and the one question a reader asks of it is whether the rows it describes fit in the bytes that follow. So a schema declaring one is a valid schema whose rows can never appear, and a document holding none of them reads.

The value types have ranges, which are what the types mean rather than limits on a document: an i is a signed 64-bit integer, an f is a finite double, and a date's year has four digits.

What bounds a document is the machine reading it: the memory it has, and the largest object its language holds at once. A reader refuses nothing else for being too large.

What a reader refuses

  • A version that is not 0.6.
  • A column with no size that is not a list column.
  • A size on a list column.
  • A size that is not a positive integer.
  • A column name containing :, a space, a newline, or NUL.
  • Two columns with the same name in one schema.
  • A type that is neither one of the nine letters nor a number.
  • A number naming a schema the document does not declare. Schemas count from 1, so 0 names none.
  • A number, boolean, date, or time that does not match its grammar.
  • A row with fewer columns than the schema. Every row has every column. There are no optional fields and no null.
  • A row count that does not match the bytes that follow.
  • A missing separator after a row.

An encoder additionally refuses a missing value, since there is no null to write in its place, a value longer than its column's size, a b value that is not exactly the size, and a break byte inside a value of any type but b.

Complete example

jalapenojson
0.6
#id:i:4 #name:s:5 #joined:d:10 #seen:z:25
2
42␀␀Alice2026-01-152026-01-15T09:30:00+02:00
7␀␀␀Bob␀␀2025-11-022025-11-02T23:05:00Z␀␀␀␀␀

Two rows of 44 bytes of field and one \n. Nothing separates the values inside a row, because each column's size says where it ends.

42 is two bytes in a four-byte field, so a break byte follows it and then one more. Alice fills its field exactly and carries none. Bob is three bytes in five, so it carries two. Z is five bytes shorter than what +02:00 needs, so Bob's seen carries five.

Alice joined on a date with no time of day and was last seen at 09:30 two hours ahead of UTC. Bob was last seen at 23:05 UTC.

examples/example.jjson is this document, byte for byte, with the real NULs in it.