1.0
0.5 · more predictable1.5 · more random

Applies to both models on the next batch.

MLP

Batch 001

A 2D embedding for each character; the previous 16 tokens pass through a 128-unit tanh layer to predict the next character.

Architecture based on Bengio et al. (2003).

  1. Starting generator…

Bigram baseline

Batch 001

A 34 × 34 count matrix from the 4,414-word training split. Each new character depends only on the previous character. View matrix

  1. Starting generator…

Grey: attested or generated in an earlier batch.