# Protobuf and reflect

> Source: https://www.ymotongpoo.com/books/tinygo-otel-esp32/20-protobuf/


## A design with the generated types

OTLP messages are defined in a [Protocol Buffers](https://protobuf.dev/) schema (protobuf from here on), and the schema lives in the [opentelemetry-proto](https://github.com/open-telemetry/opentelemetry-proto) repository. [`go.opentelemetry.io/proto/otlp`](https://pkg.go.dev/go.opentelemetry.io/proto/otlp) contains the Go types generated from that schema. You do not need the SDK to build an OTLP/HTTP request body. You can build the message with these generated types and turn it into bytes with `proto.Marshal` from [protobuf-go](https://github.com/protocolbuffers/protobuf-go).

This design cross-compiles with [TinyGo](https://tinygo.org/). The tests that run on the host also pass. When the same code sends from the host to the Collector, the Collector decodes the message correctly. I tested with `go.opentelemetry.io/proto/otlp` v1.11.0 and `google.golang.org/protobuf` v1.36.11.

## A panic on the real device

When I flash the board, though, the program stops at the first export, right after it connects to Wi-Fi.

```
wifi connected, address: 192.168.0.102
warning: clock never synchronised; timestamps will be wrong
panic: unimplemented: (reflect.Type).MethodByName()
abort called
```

The first time protobuf-go marshals a message type, it inspects the type with `reflect` and builds an internal table. The function that does this work is `makeStructInfo` (`internal/impl/message.go`). For compatibility with old generated code, it always looks up the following two methods.

```go
	if m, ok := reflect.PtrTo(t).MethodByName("XXX_OneofFuncs"); ok {
		methods = append(methods, m)
	}
	if m, ok := reflect.PtrTo(t).MethodByName("XXX_OneofWrappers"); ok {
		methods = append(methods, m)
	}
```

TinyGo's `reflect` declares `MethodByName`, but the body is only this one line (`src/reflect/type.go` in TinyGo 0.42.0).

```go
	panic("unimplemented: (reflect.Type).MethodByName()")
```

The method is declared, so the program compiles and links. It panics only when the real device reaches `MethodByName`.

This call runs whether or not the message uses oneof (the protobuf syntax for a set of fields of which a message holds only one). It is also unrelated to a shortage of memory. **A successful build does not mean that the program works.** TinyGo's `reflect` has methods that exist only as declarations and panic when they run. The linker cannot detect this gap. You find it only when the real device actually takes that path.

## The dynamicpb path

The panic happens on the path that builds type information when protobuf-go marshals a generated Go struct. protobuf-go also has [`dynamicpb`](https://pkg.go.dev/google.golang.org/protobuf/types/dynamicpb), which handles messages without generated types. `dynamicpb` uses only the descriptor that describes the structure of a message. It looks up each field by name and sets the value.

I split the test into steps inside one binary and ran it on the real device. This is the output.

```
step 1: get the OTLP descriptor from the generated type  -> OK
step 2: dynamicpb.NewMessage(desc)                       -> OK
step 3: set the nested fields                            -> OK
step 4: proto.Marshal(dynamicMessage)                    -> OK, 18 bytes
step 5: proto.Marshal(&mpb.MetricsData{})                -> panic
```

It is safe to take only the descriptor from a generated type, and the real device can marshal a `dynamicpb` message. The program fails only at the last step, which passes the generated type itself to `proto.Marshal`.

So it looks as if you could build the OTLP payload with `dynamicpb`. I wrote a program that builds a payload with only `dynamicpb`: 8 time series of metrics and 3 resource attributes. Then I measured on the real device how much memory one build allocates. The build starts like this.

```go
	desc := (&mpb.MetricsData{}).ProtoReflect().Descriptor()
	m := dynamicpb.NewMessage(desc)

	rmF := desc.Fields().ByName("resource_metrics")
	rmL := m.Mutable(rmF).List()
	rm := rmL.NewElement().Message()
	// ...
	b, err := proto.Marshal(m)
```

One build and marshal allocated 29,277 bytes of memory. The hand-written encoder in this book allocates 0 bytes for the same work. `dynamicpb` walks the descriptor and rebuilds the message tree on every call, so you cannot avoid this allocation. A device that sends every 10 seconds would throw away about 29 KB each time, and the GC would run often. This book's implementation uses less than 3,000 bytes of short-lived allocations per export (Chapter 9 covers how I kept it that low). `dynamicpb` works, but I cannot choose it for this use.

## The protobuf wire format

This book's implementation does not use the protobuf runtime. It has a `wire` package that writes the bytes directly. I can write it by hand because the [protobuf wire format](https://protobuf.dev/programming-guides/encoding/) is simple.

A protobuf message is a sequence of bytes that lists the fields one by one. Each field starts with a tag that combines the field number and the wire type (the format of the value). The value follows the tag. The tag is an integer: the field number shifted left by 3 bits, with the wire type in the low 3 bits. Integers are written as varints. A varint is a variable-length encoding. Each byte holds 7 bits of the value in its low bits, and the top bit is set when more bytes follow. Smaller values take fewer bytes.

OTLP metrics use only three wire types. varint holds enum values and booleans. fixed64 is 8 bytes of fixed length and holds timestamps and doubles. length-delimited puts the length first and holds strings and nested messages. The `wire` package writes tags and varints like this.

```go
func (w *Buffer) tag(field, wireType int) {
	w.varint(uint64(field)<<3 | uint64(wireType))
}

func (w *Buffer) varint(v uint64) {
	for v >= 0x80 {
		w.b = append(w.b, byte(v)|0x80)
		v >>= 7
	}
	w.b = append(w.b, byte(v))
}
```

I copied the field numbers from opentelemetry-proto v1.11.0 and gave them names as constants. The package uses neither `reflect` nor generated code, so the parts of `reflect` that TinyGo does not implement do not affect it.

## Reserving space for nested lengths

When you write protobuf by hand, nested messages need the most care. A length-delimited field puts the length of its contents first, but you do not know the length until you finish writing the contents. To compute the length first, you must walk the message tree once to find the size, and walk it again to write it.

`Nested` in `wire` does not walk the tree twice. It reserves the space for the length first. After it writes the contents, it computes the length. If the varint for that length is shorter than the reserved space, it moves the contents forward to close the gap.

```go
const maxLenPlaceholder = 5

// Nested writes a nested message field. The callback appends the contents.
func (w *Buffer) Nested(field int, fn func()) {
	w.tag(field, wireBytes)
	start := len(w.b)
	w.b = append(w.b, 0, 0, 0, 0, 0)
	fn()
	size := len(w.b) - start - maxLenPlaceholder

	n := varintLen(uint64(size))
	if n < maxLenPlaceholder {
		copy(w.b[start+n:], w.b[start+maxLenPlaceholder:])
		w.b = w.b[:len(w.b)-(maxLenPlaceholder-n)]
	}
	putVarint(w.b[start:start+n], uint64(size))
}
```

The payload is small (1,379 bytes for 10 time series). My judgment is that the copy to close the gap costs less than a second walk of the tree.

The reserved space must match the maximum length that the varint can take. A varint holds 7 bits per byte, so 3 bytes can express lengths up to 2,097,151. Suppose that you reserve 3 bytes. When the contents reach 2,097,152 bytes, the varint for the length becomes 4 bytes and overflows the reserved space. The code that moves the contents then breaks its own assumption, and the last byte of the contents is lost. No error comes back, so the program sends the bytes with the end missing[^placeholder]. With 5 bytes, every length that a uint32 can express fits, and the copy always moves the contents forward.

The test that checks this boundary fills the contents with position-dependent bytes.

```go
		// Position-dependent content. All-zero filler is invariant under an
		// off-by-one shift, so it would hide exactly the bug this checks.
		content := make([]byte, size)
		for i := range content {
			content[i] = byte(i%251 + 1)
		}
```

If the contents are all zeros, a shift by one byte looks like the same contents, so the test cannot detect damage from the shift. The test lines up sizes on both sides of each point where the varint for the length grows: from 1 byte to 2, from 2 to 3, and from 3 to 4. For each size, it compares the declared length and the contents, byte by byte.

## Empty strings and oneof

In proto3, the rule is that you do not send a field whose value is the zero value. The receiver reads a field that did not arrive as the zero value. The write functions in `wire` also write nothing when they receive an empty string or 0.

This rule has an exception. The OTLP attribute value `AnyValue` is a oneof, and holds one of a string, an integer, a double, and so on. In a oneof, the field that arrives decides the type of the value. If you omit an attribute value that is an empty string, the receiver gets "an attribute with no value type set". The receiver cannot tell it apart from "an attribute whose value is an empty string". For this reason, `wire` has `StringAlways`, which writes the string even when it is empty, and uses it to write attribute values.

```go
func (r *Registry) encodeAttr(a *Attr, field int) {
	w := r.buf
	w.Nested(field, func() {
		w.String(wire.FieldKeyValueKey, a.Key)
		w.Nested(wire.FieldKeyValueValue, func() {
			// AnyValue is a oneof: an empty string still has to be written,
			// or the value reads as unset instead of empty.
			w.StringAlways(wire.FieldAnyValueStringValue, a.Value)
		})
	})
}
```

For the same reason, `Double`, which writes the value of a data point, does not omit 0. A measured value of 0 has meaning, and in OTLP this value is also inside a oneof.

## A round trip through the official implementation

You cannot check on the device whether the hand-written encoder is correct, because protobuf-go does not run on the device. So a test that runs on the host reads the hand-written bytes into the generated types with the official `proto.Unmarshal`. Then it compares the field values with the expected values.

```go
func decode(t *testing.T, b []byte) *mpb.MetricsData {
	t.Helper()
	var md mpb.MetricsData
	if err := proto.Unmarshal(b, &md); err != nil {
		t.Fatalf("reference protobuf runtime rejected our bytes: %v", err)
	}
	return &md
}
```

The generated types do not run on the device, but on the host they can verify the hand-written encoder. The tests check the resource attributes, the metric names and units, the aggregation types, the values, and the timestamps one by one. They also confirm that the official implementation can read a payload larger than 2 MB.

`proto.Marshal` on the generated types passed the build and the host tests, and failed only on the real device. The build result does not decide which parts work. The result on the real device, when it takes that path, decides. The next chapter looks at the other OTLP encoding, which is separate from the protobuf binary format.

[^placeholder]: If the varint for the length is one byte longer than the reservation, the `copy` that moves the contents has a destination one byte shorter than its source. `copy` uses the shorter of the two, so it does not copy the last byte of the contents. The reslice after that extends past the length into the spare capacity. Unrelated bytes there become the end of the contents.

