# The HTTP Client and Failed Exports

> Source: https://www.ymotongpoo.com/books/tinygo-otel-esp32/30-http/


To send data over the HTTP version of [OTLP](https://opentelemetry.io/docs/specs/otlp/) (OTLP/HTTP), a client sends one HTTP/1.1 POST and reads the response. In Go, the usual choice is the standard library's [`net/http`](https://pkg.go.dev/net/http). On this book's device, I did not use `net/http` and wrote the HTTP client by hand. This chapter first checks the difference in binary size that led to that decision. It then covers the requirements that a hand-written client must meet, and what it retries and does not retry when an export fails.

## What net/http depends on

TinyGo removes unreachable code at build time. Even so, the build with `net/http` and the build with the hand-written client differ widely in flash usage. The following table shows values for the whole firmware, built with `-stack-size=16KB`. This firmware uses `encoding/json` for the metrics encoder, and it includes the Wi-Fi connection, clock synchronization, and the export of the three signals[^size-measure].

| Item | `net/http` | Hand-written | Difference |
| --- | ---: | ---: | ---: |
| Flash (bytes) | 1,067,175 | 748,691 | 318,484 |
| Static RAM (bytes) | 147,636 | 114,796 | 32,840 |

When I drop `net/http`, flash usage falls by about 320,000 bytes (30%)[^size-protobuf]. With `net/http`, `crypto/tls` enters the dependencies, and `crypto/x509`, `encoding/asn1`, and `math/big` come with it. The device speaks only plain HTTP, so these 320,000 bytes are code that never runs.

I did not test TLS communication on the device for this book. As Chapter 4 showed, TinyGo's `crypto/tls` leaves some functions unimplemented. The Wi-Fi driver [espradio](https://github.com/tinygo-org/espradio) v0.3.0 does not provide a TLS client either. So in this book's setup, the device sends plain HTTP to a Collector on the same LAN, and the Collector handles TLS to the outside. Chapter 11 explains this division of work again.

[^size-measure]: When you measure the size, you must fill in both the Wi-Fi SSID and the export endpoint with `-ldflags -X`. If either one is empty, the firmware stops with an error right after boot, by design. TinyGo then judges the network code or the export loop unreachable, removes it, and reports about one sixth of the real size.
[^size-protobuf]: With the hand-written protobuf encoder, the difference is almost the same: flash usage falls from 1,067,327 bytes to 748,907 bytes.

## The hand-written HTTP client

The hand-written client only writes the request bytes to a TCP connection and reads the response. It builds the request in a reused buffer, so that it does not allocate again for each export.

```go
r := e.req[:0]
r = append(r, "POST "...)
r = append(r, path...)
r = append(r, " HTTP/1.1\r\nHost: "...)
r = append(r, e.host...)
r = append(r, "\r\nContent-Type: "...)
r = append(r, contentType...)
r = append(r, "\r\nContent-Length: "...)
r = strconv.AppendInt(r, int64(len(body)), 10)
r = append(r, "\r\n\r\n"...)
r = append(r, body...)
```

That is all the writing side does. The reading side takes more work, and the reason is connection reuse. [lneto](https://github.com/soypat/lneto), the TCP/IP stack that [espradio](https://github.com/tinygo-org/espradio) uses, allocates send and receive buffers each time it opens a connection. If the client opens a new connection for every export, the amount of short-lived allocations on the heap grows. Chapter 9 covers that measurement. Here I look at the requirements that the client must meet once it reuses the connection.

The first requirement is to read the response to the end. An HTTP/1.1 response starts with a status line and headers that end with an empty line (two CRLFs in a row). A body follows, with the length that `Content-Length` gives ([RFC 9112](https://www.rfc-editor.org/rfc/rfc9112)). TCP can split this response at any position, so you cannot call `Read` once and parse only what arrived. If the client leaves part of the body unread, the next export reads the rest of the previous body as the status line. So the client keeps reading. It reads the headers until it finds two consecutive CRLFs, and the body until it reaches the byte count in `Content-Length`.

```go
for headerEnd < 0 {
	if len(e.resp) == cap(e.resp) {
		// Refusing is safer than parsing a truncated header block.
		e.closeConn()
		return out, errors.New("otlpmini: response header exceeds " +
			strconv.Itoa(maxHeaderBytes) + " bytes")
	}
	n, err := e.conn.Read(e.resp[len(e.resp):cap(e.resp)])
	if n > 0 {
		e.resp = e.resp[:len(e.resp)+n]
		headerEnd = indexCRLFCRLF(e.resp)
	}
	// ...
}
```

The second requirement is a limit on the header size. The HTTP specification sets no limit on the size of headers. A client that keeps reading without a limit can use any amount of heap. This book's client sets the limit at 2 KB. The headers in a Collector response are a few hundred bytes, so the limit leaves a margin of several times that size. If the headers exceed the limit, the client does not parse the partial headers. It returns an error and closes the connection.

The third requirement is to obey the server when it says that it will close the connection. If the response has `Connection: close`, the client closes the connection after it reads that response. Both the header name and the value `close` are case-insensitive ([RFC 9110](https://www.rfc-editor.org/rfc/rfc9110)), so the comparison function also ignores case. Some responses have a body of unknown length (`Transfer-Encoding: chunked`, or no `Content-Length`). The client also closes the connection for these, because it cannot find the boundary with the next response[^chunked].

[^chunked]: I did not implement parsing of chunked bodies. Collector responses carry `Content-Length`, so the device did not need it.

The fourth requirement is to reconnect when the reused connection has dropped. A Collector restart or a disconnect from the access point (AP) makes the held connection unusable. But the client does not learn that the connection dropped until its next write or read. So the client acts only when a write or a read on a reused connection fails. It then closes the connection, reconnects once, and sends the same request.

Tests that run on the host check that the client meets these requirements. Each test starts a fake server that returns responses split at fixed positions. The test then checks that the client reads them correctly.

| Requirement | What happens without it | Test |
| --- | --- | --- |
| Keep reading a split response | The client judges a correct response as broken | Split the response in the middle of the status line |
| Read the body to the end | The next export reads the previous body as the status line | Export a second time after a response with a body |
| Header limit | The client parses part of headers that exceed the limit | Return a 4 KB header |
| `Connection: close` | The client writes the next request to a connection that the server closes | Return `CONNECTION: Close` |
| Reconnect after a dropped connection | Exports keep failing after a Collector restart | Drop the connection from the server side after the first response |

I confirmed that the last requirement works on the real device. The test restarted the Collector in the middle of a 30-minute soak run. Chapter 11 covers the result.

## What to retry and what not to retry

When an export fails, the kind of failure decides whether the client can send the same data again. The OTLP specification limits the response codes that an OTLP/HTTP client should retry to four: 429, 502, 503, and 504 ([Retryable Response Codes](https://opentelemetry.io/docs/specs/otlp/#retryable-response-codes)).

> All other `4xx` or `5xx` response status codes MUST NOT be retried.

The reason for these four lies in what the codes mean. A 400 says that the content of the request is invalid. Sending the same bytes again gives the same result, and the specification states that a client must not retry a 400. A 500 is not among the four codes either, so the client does not retry it. On the other hand, 429 and 503 mean that the server temporarily cannot accept the request. 502 and 504 mean that an intermediary failed. For these, a later attempt has a chance of success.

The hand-written client sorts response codes the same way as the table in the specification.

```go
// OTLP names 429, 502, 503 and 504 as retryable. Other 5xx are not on that
// list, so they are reported as permanent rather than retried forever.
case resp.status == 429, resp.status == 502, resp.status == 503, resp.status == 504:
	return 0, errors.Join(ErrRetryable, errStatus(resp.status))
default:
	return 0, errStatus(resp.status)
```

The client also treats transport failures as retryable. Examples are a failed connect, a failed write, and a response that breaks off partway. Without a response, the client cannot know whether the request reached the Collector.

Note that the reconnect in the previous section and a retry are separate decisions. The client resends the same request on its own only when a reused connection fails at the transport stage. If the Collector returned a response, the request arrived, even when the response is a 503. So the client does not resend the request internally. Instead, it tells the caller through `IsRetryable` whether the failure is retryable.

```go
var status errStatus
var partial *PartialSuccess
if errors.As(err, &status) || errors.As(err, &partial) {
	return n, err
}
e.closeConn()
return e.export(path, body, contentType)
```

Tests also check the retry behavior for each response code. One test uses a server that returns 400, 404, 413, or 500, and checks that the client sends the payload exactly once. Another test checks that the client reports only 429, 502, 503, and 504 as retryable failures.

On the device, the caller does not keep the payload to resend it, even after a retryable failure. It counts the failure and waits for the next export cycle, 10 seconds later[^retry-after]. In the next cycle, it sends a new payload, encoded from the values at that time. A cumulative counter sends the total since boot every time. If the export of one cycle is lost, the value in the next cycle still includes the increase in that interval. The only loss is the gauge values for that cycle.

[^retry-after]: The OTLP specification also asks the client to follow the `Retry-After` header when a 429 or 503 response carries one (SHOULD). This book's client does not parse this header, and the export interval is always 10 seconds.

## Data that goes missing despite HTTP 200

An HTTP 200 response does not always mean that the Collector accepted all the data that you sent. In OTLP, when the Collector accepts only part of a request, it still returns 200 and reports the number of rejected items in the response body ([Partial Success](https://opentelemetry.io/docs/specs/otlp/#partial-success-1)). For metrics, the body is an `ExportMetricsServiceResponse` message. Its `partial_success` field holds the number of rejected data points (`rejected_data_points`) and the reason (`error_message`).

The response body comes back in the same encoding as the request. If you send protobuf, you get protobuf back, and if you send JSON, you get OTLP/JSON back. The client checks whether the body starts with `{` to tell the two encodings apart, and it parses both. The JSON field `rejectedDataPoints` is an int64, so OTLP/JSON allows it as either a string or a number (Chapter 6). The parser accepts both forms. I wanted the protobuf build to have no dependency on `encoding/json`. So I also wrote the JSON parser by hand, as a small function that only looks for the keys that it needs.

```go
case resp.status == 200:
	// HTTP 200 does not mean every point was stored. OTLP reports rejected
	// points in the body, and treating that as success hides data loss.
	if p := parsePartialSuccess(resp.body); p != nil {
		return len(body), errors.Join(ErrPartialSuccess, p)
	}
	return len(body), nil
```

The client treats an empty body, `{}`, and an empty `partialSuccess` as complete success. A healthy Collector returns these forms. If the client judged them as failures, every export would look like a failure[^warning].

[^warning]: The specification also allows a server to send a warning: it sets only `error_message` and leaves the rejected count at 0. This book's client counts a response as a partial rejection if it carries an error message, even when the rejected count is 0.

The client counts partial rejections separately from both successes and failures. The device reports export results in a counter named `device.export.attempts`. Its `outcome` attribute distinguishes `success`, `failure`, and `partial_rejection`. If partial rejections counted as successes, the success rate would hide the missing data. If they counted as failures, you could not tell them apart from failures of the transport or the Collector.

The client does not treat a partial rejection as a retryable failure either. The specification says that a client must not retry a request that received a partial success. In practice, if you send the same payload again, the Collector rejects the same data points for the same reason. The device also sends the reason for the rejection as a log, so you can read on the dashboard what the Collector rejected.

