# @awomanindatascience on Instagram

- **Type:** Video
- **Original URL:** https://www.instagram.com/p/DTcUzcLk3Ed
- **Gondola URL:** https://gondola.cc/posts/61229634-awomanindatascience-instagram
- **Thumbnail:** https://img.gondola.cc/tr:w-,h-,fo-auto/postThumbnails/a0acfdc22c.jpg
- **Posted:** 2026-01-13T07:35:55.000+00:00
- **Account Owner:** Chinar Arora (@awomanindatascience) — https://gondola.cc/awomanindatascience

## Caption

It’s Day 14 of building a LLM from scratch ✨

Most people think LLMs are complex because of code.
They’re complex because of configuration and scale.

Today I broke down the GPT-2 config that defines how the model thinks, remembers, and attends.
GPT-2 is just a set of numbers that define scale: vocab size, context length, embedding dimension, layers, and attention heads.

Breaking down the GPT-2 (124M) configuration: 50,257-token vocabulary, 1,024-token context, 768-dimensional embeddings, 12 transformer layers with 12 attention heads, dropout 0.1, and bias-free QKV projections. Understanding these parameters is key to scaling LLMs efficiently.
#deeplearning #generativeai #womenwhocode #largelanguagemodels

## Stats

- **Views:** 72,704
- **Likes:** 4,173
- **Shares:** 0
- **Comments:** 64

## Tags

largelanguagemodels, deeplearning, womenwhocode, generativeai

---
Copyright (c) Gondola