Personal Assistant
Custom Language Model - Decoder-only Transformer
Personal Assistant
Pretrained on Fineweb-Edu and SFT on SmolTalk
Endpoint status
Chat with Acai
Response will appear here.
Model Details
Inside the model.
A decoder-only Transformer designed and trained from the ground up, from tokenization and pretraining through supervised fine-tuning and autoregressive inference.
Parameters
152.6M
Decoder Layers
18
Context Length
2,048
Vocabulary
32,768
Architecture
- Model Type
- Decoder-only Transformer
- Hidden Dimension
- 768
- Attention Heads
- 12
- Feed-forward Dimension
- 2,048
- Position Encoding
- RoPE
- RoPE Base
- 10,000
- Word Embeddings
- Tied
Training & Inference
- Framework
- PyTorch
- Pretraining Dataset
- FineWeb-Edu
- SFT Dataset
- SmolTalk
- Training Objective
- Next-token Prediction
- Optimizer
- AdamW
- Inference
- Autoregressive + KV Cache
- Deployment
- AWS SageMaker
Datasets provided by Hugging Face.