Per-Layer Embeddings (PLE) in Gemma 4 explained
Gemma 4's per-layer embeddings boost model performance without increasing compute parameters
“It's a very nice way to increase the performance of the model without actually increasing parameters because these parameters are stored on flash storage.”
Google's Gemma 4 E2B and E4B models use Per-Layer Embeddings (PLE), a technique that gives each token a distinct embedding per transformer layer rather than a single shared representation. The key architectural insight is that these embeddings are stored on flash storage and only a tiny fraction is accessed during inference, meaning the model gains representational power without increasing effective compute parameters. This is a noteworthy efficiency technique relevant to edge/on-device deployment of small language models.