# Speaking OTLP/JSON

> Source: https://www.ymotongpoo.com/books/tinygo-otel-esp32/25-otlp-json/


## JSON as an encoding

The previous chapter covered a hand-written encoder that writes OTLP messages in the protobuf binary format. OTLP has a second way to write them. The [OTLP specification](https://opentelemetry.io/docs/specs/otlp/) defines two encodings for the OTLP/HTTP body: binary protobuf and JSON. Both follow the same protobuf schema. They differ only in how they turn the message into bytes. If the request's `Content-Type` is `application/x-protobuf`, the body is binary. If it is `application/json`, the body is JSON.

The OTLP receiver in the [OpenTelemetry Collector](https://opentelemetry.io/docs/collector/) accepts both encodings. So sending JSON is not a workaround for when protobuf is unavailable. It is OTLP itself, as the specification defines it.

The parts to write JSON are in the Go standard library. `encoding/json` in [TinyGo](https://tinygo.org/) works on the real device. You do not have to write an encoder yourself. You can speak OTLP by adding tags to structs.

## How the specification writes JSON

OTLP/JSON follows the standard protobuf [JSON mapping](https://protobuf.dev/programming-guides/json/). The JSON mapping is the set of rules that converts between protobuf messages and JSON. In protobuf-go, the `protojson` package implements it. The OTLP specification makes some changes to these rules. As the side that sends from the device, you need to watch these four points.

| Item | How OTLP/JSON writes it | Relation to the standard JSON mapping |
| --- | --- | --- |
| Object keys | lowerCamelCase field names (`timeUnixNano`) | The standard mapping also reads the original field names (`time_unix_nano`), but OTLP does not allow them |
| 64-bit integers | Decimal strings | Same as the standard rule |
| `traceId` and `spanId` | Hex strings | base64 in the standard mapping |
| Enum values | Integers | The standard mapping also allows name strings |

OTLP did not invent its own rule for 64-bit integers. It follows the standard protobuf rule. The specification notes that timestamps such as `timeUnixNano` and `startTimeUnixNano` are also `fixed64`, so you write them as strings. The reason is that an implementation that treats JSON numbers as float64 cannot hold UNIX time in nanoseconds (19 digits) exactly. The specification requires receivers to accept both numbers and strings. If the sender writes strings, though, it does not depend on how the receiver handles numbers.

The other three points are where OTLP departs from the standard rules. Trace IDs and span IDs are bytes, which the standard mapping writes in base64, but OTLP writes them as hex strings. Enum values must be integers, and you cannot use name strings such as `SPAN_KIND_SERVER`. The specification also requires receivers to ignore fields that they do not know.

## Messages written as struct tags

This book's `otlpjson` package writes the shape of each OTLP message directly as a Go struct, and sets the JSON keys with tags. A metric data point looks like this.

```go
type numberDataPoint struct {
	// Timestamps are strings: they are uint64 nanoseconds, and a JSON number
	// is a float64, which cannot represent them exactly.
	StartTimeUnixNano string     `json:"startTimeUnixNano,omitempty"`
	TimeUnixNano      string     `json:"timeUnixNano"`
	AsDouble          float64    `json:"asDouble"`
	Attributes        []keyValue `json:"attributes,omitempty"`
}
```

The struct holds timestamps as `string`, not `uint64`. Just before the export, `strconv.FormatUint` turns them into decimal strings. The keys are lowerCamelCase, and `AsDouble` has no `omitempty`. As the previous chapter showed, the value of a data point is a oneof, and you must send a value of 0 without omitting it.

`sum`, which represents a cumulative counter, holds the aggregation type as an enum value.

```go
// aggregationTemporalityCumulative is AGGREGATION_TEMPORALITY_CUMULATIVE.
const aggregationTemporalityCumulative = 2

type sum struct {
	DataPoints             []numberDataPoint `json:"dataPoints"`
	AggregationTemporality int               `json:"aggregationTemporality"`
	IsMonotonic            bool              `json:"isMonotonic"`
}
```

`AggregationTemporality` is an `int`. The encoder writes the value 2, not the name `AGGREGATION_TEMPORALITY_CUMULATIVE`. A trace span holds its IDs as hex strings.

```go
type span struct {
	// OTLP/JSON writes trace and span IDs as hex strings, not base64: this is
	// one of the places where the OTLP JSON mapping departs from protobuf's
	// canonical JSON.
	TraceID           string     `json:"traceId"`
	SpanID            string     `json:"spanId"`
	ParentSpanID      string     `json:"parentSpanId,omitempty"`
	Name              string     `json:"name"`
	Kind              int        `json:"kind"`
	// ...
}
```

`hex.EncodeToString` from `encoding/hex` turns the 16-byte trace ID and the 8-byte span ID into strings. Go's `encoding/json` writes `[]byte` in base64. If you put the IDs into the struct as bytes, the program sends them in a form that does not match the specification.

For metrics, logs, and traces alike, the only code I write is these structs and the code that fills in the values. I leave the job of writing the bytes to `encoding/json`.

## Checking with protojson

As with the hand-written protobuf, host tests check that the JSON is correct. The reference is `protojson`, which implements the protobuf JSON mapping. The test reads the JSON that the encoder outputs into the generated types with `protojson.Unmarshal`.

```go
func decode(t *testing.T, b []byte) *mpb.MetricsData {
	t.Helper()
	var md mpb.MetricsData
	// DiscardUnknown stays false: an unknown or misspelled field must fail
	// here rather than be silently dropped by the collector.
	if err := protojson.Unmarshal(b, &md); err != nil {
		t.Fatalf("protojson rejected our payload: %v\npayload: %s", err, b)
	}
	return &md
}
```

By default, `protojson.Unmarshal` returns an error for a field that it does not know. The OTLP specification, on the other hand, requires receivers to ignore unknown fields. So if you misspell a key, the Collector returns no error and only discards that field. If the test rejects unknown fields, you can find an error such as writing `timeUnixNanos` for `timeUnixNano` before it reaches the Collector. A separate test checks that timestamps come out as strings, not numbers.

For trace IDs only, `protojson` cannot serve as the reference. `protojson` follows the standard mapping and reads IDs as base64. Hex strings fall inside the base64 character range, so `protojson` reports no error and reads them as different bytes. So the trace test rewrites the IDs to base64 before it passes them to `protojson`, and checks the structure. A separate test checks that the IDs come out in hex. To confirm that the Collector reads hex IDs correctly, I sent traces to a real Collector and to [Grafana Tempo](https://grafana.com/oss/tempo/), the trace backend.

## Payload size

The cost of JSON is a larger payload. JSON writes every key as a string each time, and numbers also become decimal strings. So the same content is longer than in binary protobuf. I compared the payload for an export of 10 time series of metrics, using the serial output from the real device.

| Encoding | Bytes |
| --- | --- |
| Binary protobuf | 1,379 |
| JSON | 3,016 (2.19 times) |

This book's demo sends metrics every 10 seconds. Even with JSON, that is about 300 bytes per second. When the device sends to a Collector on the same LAN, this difference does not cause a problem. A battery-powered device, or a device that sends over a narrow link such as LPWAN (low-power wide-area network), has a reason to choose protobuf.

The choice of encoder barely changes the binary size. With the hand-written HTTP client, the flash difference is 216 bytes, less than 0.1% of the total[^size].

## Why JSON is the default

This book's demo uses `encoding/json` as the default encoder. A build tag switches it to the hand-written protobuf. I made JSON the default because it is the combination with the least code that I write myself[^tags].

With the hand-written protobuf, I had to copy the field numbers and write the wire format. I also had to implement nested lengths and oneof correctly myself. With JSON, I copy the shape of the OTLP messages into structs, and the standard library writes the rest. Logs and traces always use JSON, whatever the encoding for metrics is. OTLP/HTTP lets you choose the `Content-Type` per request, so I have no reason to add a hand-written encoder for each signal.

Apart from a payload a little over twice as large, choosing JSON costs nothing for this use. I kept the hand-written encoder behind the same interface for the case where a bandwidth-limited environment needs protobuf.

## encoding/json and the stack

`encoding/json` comes with one condition. It walks nested messages recursively to write them, so it uses a lot of goroutine stack. OTLP metrics are messages nested many levels deep. `scopeMetrics` sits inside `resourceMetrics`, `metrics` sits inside that, and `dataPoints` sits under `gauge` or `sum`.

The default stack for the ESP32-S3 target is 8 KB, and with that size the program crashes on the real device. This book's demo raises the stack to 16 KB at build time. Passing one small struct to `json.Marshal` does not reproduce the crash. It crashes only when you pass a whole OTLP message. Like `MethodByName` in the previous chapter, this problem stayed hidden until a real payload went through the real device. Chapter 9 shows my measurements of how the stack size relates to the crash.

[^size]: I measured the binary size of the whole device program, including the Wi-Fi connection, clock synchronization, and OTLP export, built with a 16 KB stack. The combination that sends metrics with the hand-written protobuf is 748,907 bytes. The combination that sends them with `encoding/json` is 748,691 bytes. Both send logs and traces with `encoding/json`, so the difference is only the metrics encoder.
[^tags]: Build tags select the encoder and the HTTP client independently. The default is `socket otlpjson`, which combines `encoding/json` and the hand-written HTTP client. Chapter 7 covers the HTTP client.

