First published: Last updated:

Speaking OTLP/JSON

Originally published in Japanese at https://zenn.dev/ymotongpoo/books/tinygo-otel-esp32/viewer/25-otlp-json.

JSON as an encoding

The previous chapter covered a hand-written encoder that writes OTLP messages in the protobuf binary format. OTLP has a second way to write them. The OTLP specification defines two encodings for the OTLP/HTTP body: binary protobuf and JSON. Both follow the same protobuf schema. They differ only in how they turn the message into bytes. If the request’s Content-Type is application/x-protobuf, the body is binary. If it is application/json, the body is JSON.

The OTLP receiver in the OpenTelemetry Collector accepts both encodings. So sending JSON is not a workaround for when protobuf is unavailable. It is OTLP itself, as the specification defines it.

The parts to write JSON are in the Go standard library. encoding/json in TinyGo works on the real device. You do not have to write an encoder yourself. You can speak OTLP by adding tags to structs.

How the specification writes JSON

OTLP/JSON follows the standard protobuf JSON mapping. The JSON mapping is the set of rules that converts between protobuf messages and JSON. In protobuf-go, the protojson package implements it. The OTLP specification makes some changes to these rules. As the side that sends from the device, you need to watch these four points.

ItemHow OTLP/JSON writes itRelation to the standard JSON mapping
Object keyslowerCamelCase field names (timeUnixNano)The standard mapping also reads the original field names (time_unix_nano), but OTLP does not allow them
64-bit integersDecimal stringsSame as the standard rule
traceId and spanIdHex stringsbase64 in the standard mapping
Enum valuesIntegersThe standard mapping also allows name strings

OTLP did not invent its own rule for 64-bit integers. It follows the standard protobuf rule. The specification notes that timestamps such as timeUnixNano and startTimeUnixNano are also fixed64, so you write them as strings. The reason is that an implementation that treats JSON numbers as float64 cannot hold UNIX time in nanoseconds (19 digits) exactly. The specification requires receivers to accept both numbers and strings. If the sender writes strings, though, it does not depend on how the receiver handles numbers.

The other three points are where OTLP departs from the standard rules. Trace IDs and span IDs are bytes, which the standard mapping writes in base64, but OTLP writes them as hex strings. Enum values must be integers, and you cannot use name strings such as SPAN_KIND_SERVER. The specification also requires receivers to ignore fields that they do not know.

Messages written as struct tags

This book’s otlpjson package writes the shape of each OTLP message directly as a Go struct, and sets the JSON keys with tags. A metric data point looks like this.

type numberDataPoint struct {
	// Timestamps are strings: they are uint64 nanoseconds, and a JSON number
	// is a float64, which cannot represent them exactly.
	StartTimeUnixNano string     `json:"startTimeUnixNano,omitempty"`
	TimeUnixNano      string     `json:"timeUnixNano"`
	AsDouble          float64    `json:"asDouble"`
	Attributes        []keyValue `json:"attributes,omitempty"`
}

The struct holds timestamps as string, not uint64. Just before the export, strconv.FormatUint turns them into decimal strings. The keys are lowerCamelCase, and AsDouble has no omitempty. As the previous chapter showed, the value of a data point is a oneof, and you must send a value of 0 without omitting it.

sum, which represents a cumulative counter, holds the aggregation type as an enum value.

// aggregationTemporalityCumulative is AGGREGATION_TEMPORALITY_CUMULATIVE.
const aggregationTemporalityCumulative = 2

type sum struct {
	DataPoints             []numberDataPoint `json:"dataPoints"`
	AggregationTemporality int               `json:"aggregationTemporality"`
	IsMonotonic            bool              `json:"isMonotonic"`
}

AggregationTemporality is an int. The encoder writes the value 2, not the name AGGREGATION_TEMPORALITY_CUMULATIVE. A trace span holds its IDs as hex strings.

type span struct {
	// OTLP/JSON writes trace and span IDs as hex strings, not base64: this is
	// one of the places where the OTLP JSON mapping departs from protobuf's
	// canonical JSON.
	TraceID           string     `json:"traceId"`
	SpanID            string     `json:"spanId"`
	ParentSpanID      string     `json:"parentSpanId,omitempty"`
	Name              string     `json:"name"`
	Kind              int        `json:"kind"`
	// ...
}

hex.EncodeToString from encoding/hex turns the 16-byte trace ID and the 8-byte span ID into strings. Go’s encoding/json writes []byte in base64. If you put the IDs into the struct as bytes, the program sends them in a form that does not match the specification.

For metrics, logs, and traces alike, the only code I write is these structs and the code that fills in the values. I leave the job of writing the bytes to encoding/json.

Checking with protojson

As with the hand-written protobuf, host tests check that the JSON is correct. The reference is protojson, which implements the protobuf JSON mapping. The test reads the JSON that the encoder outputs into the generated types with protojson.Unmarshal.

func decode(t *testing.T, b []byte) *mpb.MetricsData {
	t.Helper()
	var md mpb.MetricsData
	// DiscardUnknown stays false: an unknown or misspelled field must fail
	// here rather than be silently dropped by the collector.
	if err := protojson.Unmarshal(b, &md); err != nil {
		t.Fatalf("protojson rejected our payload: %v\npayload: %s", err, b)
	}
	return &md
}

By default, protojson.Unmarshal returns an error for a field that it does not know. The OTLP specification, on the other hand, requires receivers to ignore unknown fields. So if you misspell a key, the Collector returns no error and only discards that field. If the test rejects unknown fields, you can find an error such as writing timeUnixNanos for timeUnixNano before it reaches the Collector. A separate test checks that timestamps come out as strings, not numbers.

For trace IDs only, protojson cannot serve as the reference. protojson follows the standard mapping and reads IDs as base64. Hex strings fall inside the base64 character range, so protojson reports no error and reads them as different bytes. So the trace test rewrites the IDs to base64 before it passes them to protojson, and checks the structure. A separate test checks that the IDs come out in hex. To confirm that the Collector reads hex IDs correctly, I sent traces to a real Collector and to Grafana Tempo, the trace backend.

Payload size

The cost of JSON is a larger payload. JSON writes every key as a string each time, and numbers also become decimal strings. So the same content is longer than in binary protobuf. I compared the payload for an export of 10 time series of metrics, using the serial output from the real device.

EncodingBytes
Binary protobuf1,379
JSON3,016 (2.19 times)

This book’s demo sends metrics every 10 seconds. Even with JSON, that is about 300 bytes per second. When the device sends to a Collector on the same LAN, this difference does not cause a problem. A battery-powered device, or a device that sends over a narrow link such as LPWAN (low-power wide-area network), has a reason to choose protobuf.

The choice of encoder barely changes the binary size. With the hand-written HTTP client, the flash difference is 216 bytes, less than 0.1% of the total1.

Why JSON is the default

This book’s demo uses encoding/json as the default encoder. A build tag switches it to the hand-written protobuf. I made JSON the default because it is the combination with the least code that I write myself2.

With the hand-written protobuf, I had to copy the field numbers and write the wire format. I also had to implement nested lengths and oneof correctly myself. With JSON, I copy the shape of the OTLP messages into structs, and the standard library writes the rest. Logs and traces always use JSON, whatever the encoding for metrics is. OTLP/HTTP lets you choose the Content-Type per request, so I have no reason to add a hand-written encoder for each signal.

Apart from a payload a little over twice as large, choosing JSON costs nothing for this use. I kept the hand-written encoder behind the same interface for the case where a bandwidth-limited environment needs protobuf.

encoding/json and the stack

encoding/json comes with one condition. It walks nested messages recursively to write them, so it uses a lot of goroutine stack. OTLP metrics are messages nested many levels deep. scopeMetrics sits inside resourceMetrics, metrics sits inside that, and dataPoints sits under gauge or sum.

The default stack for the ESP32-S3 target is 8 KB, and with that size the program crashes on the real device. This book’s demo raises the stack to 16 KB at build time. Passing one small struct to json.Marshal does not reproduce the crash. It crashes only when you pass a whole OTLP message. Like MethodByName in the previous chapter, this problem stayed hidden until a real payload went through the real device. Chapter 9 shows my measurements of how the stack size relates to the crash.


  1. I measured the binary size of the whole device program, including the Wi-Fi connection, clock synchronization, and OTLP export, built with a 16 KB stack. The combination that sends metrics with the hand-written protobuf is 748,907 bytes. The combination that sends them with encoding/json is 748,691 bytes. Both send logs and traces with encoding/json, so the difference is only the metrics encoder. ↩︎

  2. Build tags select the encoder and the HTTP client independently. The default is socket otlpjson, which combines encoding/json and the hand-written HTTP client. Chapter 7 covers the HTTP client. ↩︎