# Memory, GC, and Stacks

> Source: https://www.ymotongpoo.com/books/tinygo-otel-esp32/40-memory/


The internal SRAM of the ESP32-S3 is 416 KB, at least four orders of magnitude less than the memory of a server. So when you send telemetry from a microcontroller, your first worry is the amount of memory. When I measured it, this book's device had enough memory. The problems were the amount of heap that each export allocates and throws away, and the size of goroutine stacks. Using values that I measured on the real device, this chapter looks at the heap, then the region that the Wi-Fi driver uses, then the stack.

## The amount of memory

The DRAM that TinyGo can use is 425,984 bytes (416 KB), and the heap limit within it is 298,287 bytes. Static regions, such as global variables, and the main stack take the rest. The real device reported the following heap values itself.

| Item | Bytes | Share of DRAM |
| --- | ---: | ---: |
| DRAM that TinyGo can use | 425,984 | 100% |
| Heap limit | 298,287 | 70% |
| Heap after GC (metrics only) | about 124,000–179,000 | 29–42% |
| Heap after GC (three signals) | about 204,000–206,000 | 48% |

The heap that remains right after a GC corresponds to the amount of data that the program actually holds. Even in the current version, which sends three signals (metrics, logs, and traces), this value is less than half of the DRAM. The heap still has a margin of about 90,000 bytes before it reaches its limit.

A time series of heap usage, however, does not look like it has a margin. Usage grows with each export, and when it gets close to the limit, a GC runs and usage drops. This sawtooth shape repeats. TinyGo's GC is a conservative GC, which treats every value that looks like a pointer as a pointer. It starts to collect when the heap has no free space and an allocation fails. So the top of each tooth sticks to the heap limit, and if you look only at the tops, memory looks insufficient. The value that shows what the program holds is the bottom of each tooth.

![Heap usage over time in Grafana](tinygo-otel-grafana-dashboard.png)
*Figure 1: The "Go heap in use" panel in the center of the top row. Usage grows as exports repeat and drops at each GC, in a sawtooth pattern.*

## Short-lived allocations

The slope of the sawtooth depends on the number of bytes that one export allocates on the heap and then throws away. This book calls this amount **churn** (short-lived allocations). The larger the churn, the shorter the interval between GCs, and each GC pause disturbs the export cycle by the length of the pause.

In the first implementation, the heap grew by 8,272 bytes per export. The value was exactly the same every time, so I suspected a leak. But when a GC ran, the heap went back to its earlier level, so all of the growth was discarded data. The 8,272 bytes came from two sources.

The first source was that the device opened a new TCP connection for every export. [lneto](https://github.com/soypat/lneto) is the TCP/IP stack that the Wi-Fi driver [espradio](https://github.com/tinygo-org/espradio) uses. Each time lneto opens a connection, it allocates a 4,096-byte transmit buffer (TxBuf) and a 1,024-byte receive buffer (RxBuf). These buffers, together with queues and a connection wrapper, were thrown away at every export. When I kept the connection open and reused it, the growth per export dropped to 4,816 bytes. To reuse a connection, the client must read the HTTP response to the end. Chapter 7 covers that requirement.

The second source was `json.Marshal` in [`encoding/json`](https://pkg.go.dev/encoding/json). `json.Marshal` encodes into an internal buffer, then copies the result into a new slice and returns it.

```go
buf := append([]byte(nil), e.Bytes()...)
```

Because the API returns the result as a slice, every call must allocate a new region. So I switched to `json.Encoder`, which writes to an `io.Writer`. I created a single `bytes.Buffer` as its destination and reused it.

```go
// json.Marshal ends with append([]byte(nil), ...), so it allocates a new
// buffer on every call. json.Encoder writes into an io.Writer instead,
// which lets one buffer be reused for the life of the program. Measured
// on M5Stack CoreS3: this removes 4,816 bytes of heap churn per export.
buf *bytes.Buffer
enc *json.Encoder
```

The `Encode` method of `json.Encoder` appends a newline at the end. The code trims that last newline before it returns, so that the bytes that the device sends match what the Collector interprets. With this change, the growth per export became 688 bytes.

After that, I added code to read the response to the end and to parse the partial success response (Chapter 7), and code to export logs and traces (Chapter 10). So the churn of the current version has grown again.

| Implementation | Heap growth per export (bytes) |
| --- | ---: |
| Open a connection for every export and use `json.Marshal` | 8,272 |
| Reuse the connection | 4,816 |
| Also reuse the buffer | 688 |
| Add reading the response to the end and parsing partial success (metrics only) | 800–848 |
| Add logs and traces (current three-signal version) | 2,992 |

In the version that sends only metrics, a 30-minute soak run made 180 exports, and 176 of them grew the heap by exactly 800 bytes.[^delta] The two exceptions matched the moments when the device opened a connection. The first connection added 46,848 bytes, and the reconnection after a Collector restart added 6,944 bytes. GC ran about once every 20 minutes.

In each export cycle, the current three-signal version sends three POSTs, for metrics, logs, and traces. It also converts the IDs of four spans to hex strings and assembles the payloads. So the growth per cycle is 2,992 bytes. In a 30-minute test, all 170 cycles had this value, except the two cycles in which the device reopened the connection. GC ran 7 times in 30 minutes, about once every 5 minutes. The heap after GC was around 204 KB all 7 times, and it did not keep growing.[^protobuf-churn]

[^delta]: These counts are the heap growth between two consecutive exports. 180 exports give 179 differences. I did not count the intervals in which GC ran, because the heap shrinks there. For the metrics-only version, 179 minus 1 GC and 2 connections leaves 176. For the three-signal version, 179 minus 7 GCs and 2 reconnections leaves 170.
[^protobuf-churn]: The metrics-only version with the hand-written protobuf encoder grows the heap by 2,752 bytes per export, which is more than the JSON version. In host benchmarks, the hand-written encoder makes zero allocations. So I think that the cause is a path outside the encoder, such as response parsing, but I have not investigated it yet.

When I reduced the churn, it also mattered that the growth was the same every time, in addition to being small. If the growth is constant per export, you can calculate the interval between GCs in advance. If the value changes at some point, you know that an unusual allocation happened in that export.

## Memory that the GC cannot see

The DRAM also holds a region that the Wi-Fi driver uses, separate from the Go heap. espradio links the Wi-Fi binary (blob) that Espressif provides. The blob takes memory from a fixed-size arena (a region that it allocates in advance). This arena is outside the control of the Go GC, so it does not appear in the heap values of `runtime.ReadMemStats`.

The device sends the arena usage and capacity that espradio's `ArenaStats` returns, as metrics separate from the Go heap. With the two sent separately, you can tell whether a memory problem comes from the Go code or from the Wi-Fi driver. On the real device, the arena used 31,312 of its 49,144 bytes, and the value did not change during the soak run.

## What adding PSRAM does not change

The M5Stack CoreS3 has 8 MB of PSRAM (external RAM) in addition to the internal SRAM. However, the TinyGo 0.42.0 linker script for the ESP32-S3 covers only the 416 KB of internal SRAM, and no code uses the PSRAM. [PR #5554](https://github.com/tinygo-org/tinygo/pull/5554) is the work to add PSRAM support for the ESP32-S3.[^psram]

[^psram]: At the time of writing, PR #5554 is not merged. The PR description says that PSRAM on an OPI connection is tested, and that a QSPI connection is not tested.

I think that this book's device would behave the same even if PSRAM became usable. The data that the program holds fits in the internal SRAM. PSRAM would only increase the amount of discarded data that can pile up before a GC runs. PSRAM sits on an external bus, so it is also slower than the internal SRAM. Uses that need a large contiguous region, such as a screen frame buffer or audio, are a different case.

## Goroutine stacks

After the amount of memory, the next problem was the size of goroutine stacks. In the standard Go toolchain, when a goroutine stack runs short, the runtime copies it to a larger region to grow it. A TinyGo goroutine stack does not grow. TinyGo allocates it at a fixed size when it creates the goroutine. The `-stack-size` build option sets that size, and the default for the ESP32-S3 is 8 KB.

`encoding/json` walks Go values recursively to encode them. An OTLP message is deeply nested: `resourceMetrics` contains `scopeMetrics`, which contains `metrics`, which contains `sum` and then `dataPoints`. So the recursion is deep too. A single struct or slice that you pass to `json.Marshal` does not crash even with 8 KB. The crash appears only when you pass a whole OTLP payload, so small tests do not find it.

I changed `-stack-size` on the real device and measured the boundary (the encoder was `encoding/json`, and the HTTP client was hand-written).

| `-stack-size` | Result | Symptom |
| --- | --- | --- |
| 8 KB (default) | Crashes | `EXCCAUSE = 28 (LoadProhibited)` |
| 12 KB | Crashes | `EXCCAUSE = 28 (LoadProhibited)` |
| 14 KB | Crashes | `fatal error: goroutine stack overflow` |
| 15 KB | Runs | |
| 16 KB | Runs | Passed a 30-minute soak run |

The boundary lies between 14 KB and 15 KB. With 15 KB, this book's device has only 1 KB of margin, and a slightly larger payload could crash it again. So I use 16 KB.

Note that in this table, the symptom changes with how short the stack is. At 14 KB, TinyGo reports `goroutine stack overflow` and stops. At 12 KB or less, the device crashes with a `LoadProhibited` exception before that report appears. `LoadProhibited` is the exception that occurs when the CPU tries to read an address that it is not allowed to read. EXCCAUSE is the register that holds the cause of an exception, and its value 28 indicates this exception.

I think that the way TinyGo detects a stack overflow explains this difference. TinyGo allocates each goroutine stack from the heap. It writes a fixed value (a canary) into the word at the bottom of the stack (the lowest address). When a goroutine suspends and returns to the scheduler, TinyGo checks whether this value has changed. If it has changed, TinyGo decides that the stack overflowed and stops.[^canary] The check happens only when a goroutine suspends. If the stack overflows before then, nothing stops it, and the overflow overwrites heap memory outside the stack. If the overflow is small, execution reaches the check before anything uses the overwritten memory, and the report appears. If the overflow is large, the program reads an overwritten pointer and hits `LoadProhibited` before it reaches the check. That is my interpretation.

[^canary]: You can confirm this in `src/internal/task/task_stack.go` and `task_stack_unicore.go` in TinyGo 0.42.0. Only `task.Pause` checks the canary.

So the absence of a stack overflow error does not mean that the stack is large enough. When the device crashes with `LoadProhibited`, keep the stack size on your list of suspects.

Finally, the `-stack-size` value requires a unit. If you write `-stack-size=16384`, TinyGo rejects the build with the error `Unrecognized size suffix`. Write `-stack-size=16KB` instead.

