Why dates look like this
The README says what the four date tags are and how to use them. This file says why they are shaped that way, and what was considered and rejected, so the next person to touch them does not have to re-derive it.
What d used to be
Before the four tags, d existed but meant nothing. Python's decode routed it
to the same branch as s, the JS decoder had a literal // s, d on that line,
and the C parser stored the tag byte without ever reading it. Nothing validated
a d value and nothing returned a date, so "not a date at all" was a
perfectly good d.
Encoding was worse than untyped. Both encoders simply stringified whatever they were handed, so the same value produced different bytes depending on the language and the machine:
Python date(2026,1,15) -> 2026-01-15
Python datetime(2026,1,15,9,30) -> 2026-01-15 09:30:00 (a space, not "T")
JS new Date(Date.UTC(...)) -> Thu Jan 15 2026 09:30:00 GMT+0000 (Coordinated Universal Time)The JS spelling also moved with the writer's own timezone: the identical Date
encoded to 09:30:00 GMT+0000 on a UTC machine and 18:30:00 GMT+0900 on one
in Tokyo. Two machines writing the same value produced different files. For a
serialization format that is a portability hole, and it is the concrete bug the
four tags exist to close. js/test_jalapenojson.mjs still runs a case for it.
Why four tags and not one
The distinction that matters, and that most formats get wrong, is **wall clock versus absolute instant**:
| Concept | Example | What did the clock read? | Which instant? |
|---|---|---|---|
| Local date | 2026-01-15 | partly | no |
| Time of day | 09:30:00 | yes | no |
| Naive datetime | 2026-01-15T09:30:00 | yes | no |
| Datetime + offset | 2026-01-15T09:30:00+02:00 | yes | yes |
A naive datetime is not a point in time; it is a description of a clock face.
A datetime with an offset is both, which is why z is the one to default to.
Collapsing these into one tag is what forces the guessing that caused the
original bug — you cannot write a date without inventing a time, or an instant
without inventing a zone. Leaving n out in particular pushes people to write a
made-up offset into a z column, which is worse than saying "no zone here".
The tags cost nothing per row, because the type is in the header once per column rather than beside each value. That is why widening from one tag to four was cheap enough to be worth doing properly.
Why ASCII and not fixed-width binary
The obvious alternative is binary: int32 days since the epoch, int64 microseconds, int16 offset minutes. That is 4–10 bytes instead of 10–36, sorts as integers, and needs no parsing at all — which sounds like exactly this format's ethos.
It was rejected because jalapenojson stores integers and floats as ASCII. A
binary date would be the first binary numeric in the format, and it would be odd
for dates to be better optimized than i. Going binary is a coherent design,
but it is a decision about the whole format rather than about dates, and it
costs the ability to read a document by eye. If it is ever revisited, note that
a binary tag would have to be routed alongside b, which is the one type whose
bytes are handed back untouched rather than UTF-8 decoded.
Why there is no zoned datetime
A tag carrying an IANA name (2026-01-15T09:30:00+02:00[Europe/Oslo]) was
considered and left out. It is only needed to keep a future local time correct
across a DST rule change — if a government moves the changeover, an offset
recorded today becomes wrong while a zone name stays right. That is real but
rare, and it drags a timezone database into every implementation.
Until that case shows up, z plus a separate s column holding the zone name
does the same job with no new machinery.
Why the two limits
Both exist so that no implementation can write a value another one cannot read. A format whose three implementations disagree about a value is back to the problem the tags were added to fix.
- Fractional seconds stop at 6 digits, and are written with trailing zeros trimmed so that
.123is the only spelling of 123 milliseconds. Six is the cap because a Pythondatetimehas a microsecond field and no nanosecond one:datetime(..., microsecond=123456789)raises. At six digits Python is exactly lossless; at nine it could not return a plaindatetimeat all, so the Python API would get worse for every caller in order to serve a rare case. That is the real argument for the cap, and it is worth knowing before anyone relaxes it — the question is not "can C carry nanoseconds" (it can) but "what does Python hand back".
A JavaScript Date is coarser than either, holding milliseconds only, which
is why the JS decoder returns micros alongside the Date rather than
putting the fraction inside it. An earlier version did not, and silently
truncated .123456 to .123 on a decode-and-re-encode.
Trimming is trailing-only, which is worth stating because it is easy to get
backwards: half a second is .5, five milliseconds is .005, and a fraction
of zero loses the dot altogether. All three implementations have that pinned
in a test.
- Second 60 is rejected. RFC 3339 allows it for leap seconds, but neither a Python
timenor a JSDatecan represent one. Accepting it in C alone would produce files only C could read.
Both are stricter than RFC 3339. That is deliberate, and it is the reason to think twice before relaxing either: the cost is not "some values are rejected", it is "the implementations stop agreeing".
Why validation happens on access, not at parse
The format's promise is that a parser never scans value bytes. Validating dates during parsing would break it, and would make date columns slower than every other column.
So validation lives wherever a value is already being materialized. In C that is
the accessors in jj_date.h, which is why jalapenojson.h needed no change at
all and jj_parse still reads a length and jumps — validating and converting
50,000 z values costs about 0.75 ms, and only when you ask. In Python and
JavaScript, decode() already is the materializing step (it is where i
becomes an int), so date validation sits in the same place.
What each language hands back, and why JS is the odd one
Python maps cleanly onto date, time, naive datetime and aware datetime.
C is fine because jj_date is defined here.
JavaScript has no type for a date without a zone, so d, t and n come back
as canonical strings.
z cannot come back as a bare Date, and this is a constraint on any future
reshaping of the JS decoder rather than a detail of today's one. A Date is an
instant: it carries no offset, so decoding into one alone discards the +02:00
and re-renders in the reader's own timezone — the original bug wearing a
different hat — and it holds only milliseconds, so it silently truncates
anything finer. **Both the offset and the fractional second have to be carried
beside the instant**, whatever the shape is called. Temporal would solve both
properly and is worth revisiting once it is universally available.