Last updated 8/19/25 — xperyv4 posted
Here I share my thoughts, implementations, and explanations of my LLM project. Earlier on you'll find conceptual and mathematical explanations of the entire architecture — attention mechanisms, transformer blocks, and more. Later on, evaluations and reflections from continuous training and revising of my models.
Posts
- Mathematics of Attention Mechanisms
- Implementation of Weights, Self Attention
- Dropouts, Masking, Casual Attention
- Re-implementing Attention Mechanisms
- Early Explorations of GPT Architecture
- nn.Linear Steps and Normalization
- Deeper Normalization, GELU, Feed Forward
- Shortcut Connections, Transformers, and Analytics
- "It's alive! It's alive!"
- Model Outputs, Cross Entropy, other Evaluations
- Chores and Training
- Integration, Experimentation, Temperature, Top K
- Training with Gutenberg
- Gutenberg Begins Training, 36 Hours to Go
- Exuperyv2 Out, Early Looks
- Exploring Existing Weights, Reconsidering Paths
- Reflections on v3 and v4