JSON vs YAML: When to Use Each
One question decides it: will a human type this by hand? Everything else follows from the answer.
JSON and YAML describe essentially the same data model — nested mappings, sequences and scalars — but they were designed for opposite readers. Understanding that is enough to settle almost every "which should we use" argument.
Why two formats exist
JSON emerged from JavaScript object literal syntax in the early 2000s and was standardised as RFC 8259. Its design priorities were machine parsing: unambiguous grammar, minimal syntax, fast to parse. It succeeded completely — JSON is now the default wire format for essentially every web API.
YAML predates it and was designed for a different problem: humans writing configuration by hand. Its priorities were readability and writability — indentation instead of braces, no mandatory quote characters, and comments.
So the split is not technical superiority. It is audience.
Side by side
// JSON
{
"service": "api",
"replicas": 3,
"ports": [80, 443],
"env": {
"LOG_LEVEL": "info"
}
}
# YAML
service: api
replicas: 3
ports:
- 80
- 443
env:
LOG_LEVEL: info
# raise to debug when triaging the latency issue
The YAML version is shorter and, crucially, can carry the comment explaining why a value is what it is. That comment is often the most valuable part of a configuration file, and JSON structurally cannot contain it.
Where each one wins
| JSON | YAML | |
|---|---|---|
| Designed for | Machines | Humans writing by hand |
| Comments | Not supported | Supported |
| Specification | RFC 8259, tight | Complex, multiple versions |
| Parse speed | Faster — simpler grammar | Slower — indentation resolution |
| Ambiguity | None | Implicit typing, several footguns |
| Multi-document | No | Yes, via --- |
| Reuse | No | Anchors, aliases, merge keys |
| Tooling maturity | Universal, excellent | Good, patchier |
| Error messages | Usually precise | Often cryptic |
| Typical use | APIs, storage, transport | Kubernetes, CI, IaC |
YAML's genuine traps
YAML's flexibility is not free. These are real, recurring problems rather than theoretical objections.
The Norway problem
YAML 1.1 treats y, yes, no, on,
off and ~ as booleans and nulls. A file listing country codes
breaks the moment it contains:
countries:
- no: Norway # parsed as false: Norway
- on: Ontario # parsed as true: Ontario
YAML 1.2 largely fixed this, but parser support varies and several widely used tools still exhibit 1.1 behaviour. The fix is always the same: quote anything that could be misread.
Implicit typing surprises
version: 1.10 # number → serialises as 1.1
version: "1.10" # string → stays 1.10
port: 8080 # number
port: "8080" # string — matters if your consumer expects a string
Indentation is semantic
YAML forbids tabs for indentation. Mixing tabs and spaces, or being one space off from siblings, produces either a parse error or — worse — a valid document with the wrong structure. In a 500-line file, a single misaligned key can silently change the resulting object, and the error message will usually point somewhere unhelpful.
Anchors make files hard to reason about
Anchors (&ref) and aliases (*ref) enable reuse, but a
reader has to mentally resolve them. The merge key <<: *base adds
another layer. Used sparingly they reduce duplication; used heavily they make a file
effectively impossible to read linearly.
JSON's real limits
JSON is not without drawbacks, and the main one is significant in practice.
No comments
For configuration this is genuinely painful. The workarounds are all imperfect: a
_comment key that pollutes the schema, a sibling README, or
JSONC — which is a superset that accepts comments and trailing commas but requires a
JSONC-aware parser, so it is not standard JSON and will not work everywhere.
Verbose for deeply nested data
Braces, quotes and commas accumulate. For a heavily nested configuration file, JSON is materially harder to scan than the YAML equivalent.
No trailing commas
Hand-editing JSON means constantly managing comma placement, and the error message for a misplaced comma points at the following line rather than the actual problem.
The decision rule
Ask one question: will a human type this by hand?
- Yes → YAML. Kubernetes manifests, CI pipelines, Docker Compose, Ansible, OpenAPI specs, application configuration. Comments and reduced punctuation are worth the parsing quirks. This is why the entire cloud-native ecosystem converged on YAML.
- No → JSON. API payloads, data at rest, generated files, anything transmitted over a network. The grammar is tighter, parsers are faster and more consistent, and you will never debug an indentation error at 2 a.m.
Two corollaries worth internalising:
Separate generated config from hand-written config. A common and effective pattern is a small hand-written YAML file of settings, plus generated JSON artifacts. Each format does what it is good at.
Do not fight your ecosystem. If you are writing Kubernetes manifests, use YAML — fighting it with JSON is possible but you lose comments and every example you find will be YAML. Conversely, if you are building an API, JSON is the expectation.
Converting between them
JSON → YAML is always safe and lossless: any valid JSON document is valid YAML (under YAML 1.2). YAML → JSON is lossy in specific, predictable ways:
- Comments are dropped. There is nowhere to put them.
- Anchors and aliases are expanded. You get the resolved data but lose the reuse structure.
- Multi-document files have no JSON equivalent — convert one document at a time.
- Custom tags (
!Ref,!!binary) have no representation.
From the command line:
# YAML → JSON
python3 -c "import json,sys,yaml; json.dump(yaml.safe_load(open(sys.argv[1])), sys.stdout, indent=2)" config.yaml
# JSON → YAML
python3 -c "import json,sys,yaml; yaml.safe_dump(json.load(open(sys.argv[1])), sys.stdout)" config.json
Note safe_load. Plain yaml.load can construct arbitrary
Python objects and is a remote code execution risk on untrusted input. Always use
safe_load / safe_dump unless you specifically need rich tags
and you trust the source completely.
For quick one-off conversions without installing anything, use our JSON ⇄ YAML Converter, which runs entirely in your browser.
Frequently asked questions
Under YAML 1.2, yes — JSON is effectively a subset. Under YAML 1.1 there are edge cases. The reverse never holds: YAML has comments, anchors and multi-document support that JSON cannot represent.
Unquoted 1.10 parses as a number and trailing zeros are not significant numerically. Quote it to keep it a string.
JSON, typically by a meaningful margin. Its grammar is simpler and requires no indentation resolution. For configuration read once at startup the difference is irrelevant; for high-throughput APIs it matters.
Not in standard JSON. Options are JSONC (requires a compatible parser), a _comment key, or keeping documentation in a separate file. If comments are essential, YAML is the better format.
Not for untrusted input — yaml.load can instantiate arbitrary Python objects, which is a code execution risk. Always use yaml.safe_load unless you have a specific reason not to.
Related guides
How AI Token Counting Works
A token is not a word, and the difference between the two is where most LLM cost estimates go wrong. Here is how tokenizers actually split your text.
LLM API Cost Compared
Published per-million-token prices are only half the equation. Tokenizer differences, caching, batching and reasoning tokens change the real number.
JWT Explained
A JWT is three Base64 segments and a signature. Understanding what the signature does — and does not — prevent is the whole game.