How a small model gets made
No distillation, no borrowed checkpoints. Data, tokenizer, architecture and training loop are our own code — small enough to iterate on in a day.
Our first model
A 124M parameter conversational model, pretrained from scratch and refined with supervised fine-tuning to follow instructions, hold natural conversations and keep a consistent identity.
- Own data pipeline and tokenizer — nothing borrowed
- Trainable on a single consumer GPU
- Benchmarked zero-shot against other ~125M models