Andrej Karpathy
d05e848010
fix pytorch code bug forward needs to know about backward_grad_sync
2024-06-17 19:49:36 +00:00
Andrej
077451927f
Merge pull request #546 from kmyusk/master
...
Fix periodic inference during training
2024-06-07 21:37:51 -07:00
Aleksa Gordic
6f20ee9f8e
Remove * 2 for total batch size
2024-06-04 23:42:54 +02:00
Aleksa Gordic
34766253cf
Move all reduce outside of the gradacc loop for efficiency
2024-06-04 23:39:52 +02:00
Aleksa Gordic
724085ecd6
tmp
2024-06-04 23:23:11 +02:00
Aleksa Gordic
3e44d275d8
Update comment
2024-06-04 18:48:29 +02:00
Aleksa Gordic
ec92ff64a2
Improve the design and comment
2024-06-04 18:47:23 +02:00
Aleksa Gordic
54ec686649
Fix zero grads bug
2024-06-04 18:40:54 +02:00
Yusik Kim
1ee0b437ca
Fix periodic inference during training
...
Problem: Enabling the generation block crashes the program
Solution: Change the variable names for x and y so that it does not
affect downstream code
Testing: Runs as expected with tinyshakespeare
Signed-off-by: Yusik Kim <kmyusk@gmail.com >
2024-06-04 10:28:55 +02:00
Andrej Karpathy
5f547601a0
further hide the c pytorch bridge logic
2024-05-30 18:20:29 +00:00
Andrej Karpathy
ca27eb8fa5
bugfix don't train on val batch
2024-05-30 18:18:05 +00:00
Andrej Karpathy
9b935a8f04
fix bug with advancing position in dataloader
2024-05-30 13:16:07 +00:00
Andrej Karpathy
555d07502a
more feature parity work, i think it's quite close now
2024-05-30 02:29:40 +00:00
Andrej Karpathy
03c2e85601
sharded dataloader for PyTorch
2024-05-30 02:05:06 +00:00
Andrej Karpathy
3fdf88e4d4
add logfile similar to C
2024-05-29 23:59:35 +00:00
Andrej Karpathy
083c50a52e
small stylistic touchups that match the c code
2024-05-29 23:50:13 +00:00
Andrej Karpathy
c413bfe0ca
pytorch parity v1
2024-05-29 23:25:55 +00:00
Chinthaka Gamanayakege
f2809a6bdb
zero changes for torch
2024-05-27 06:44:54 +00:00
Andrej Karpathy
9f08882051
add weight decay, but only for 2D tensors, as done in GPT series and in general too. this forces us to break up our adamw kernel again into one call per tensor, so there is a small throughput hit, of about 0.5% for me. but we have to break up this kernel in near future anyway
2024-05-25 00:14:10 +00:00
Andrej Karpathy
032e76c259
start putting llm.c and pytorch right next to each other, identical training runs with identical results and prints. almost
2024-05-24 01:02:21 +00:00
Andrej Karpathy
645869b6f7
add weight decay -c option and be more careful in our tests of correctness, our weight decay didn't match to pytorch. also modify the betas in AdamW to be consistent with those used in GPT-3 training
2024-05-23 20:45:59 +00:00
Andrej Karpathy
86682af9a9
llm.c matches pytorch init from scratch exactly now
2024-05-23 17:50:33 +00:00
Andrej Karpathy
e6a7d1d3e9
allow the python script to also init from random and save those weights, so it's a good reference for our C implementation
2024-05-23 14:59:25 +00:00
Andrej Karpathy
587506d096
torch tensor can't handle uint16 so let's convert to int32, which is silly because we'll convert to .long right after but ok
2024-05-21 19:22:09 +00:00
Andrej Karpathy
7d58fd2abb
adjust py file as well and make the errors better
2024-05-21 18:29:55 +00:00
Andrej Karpathy
9bad49a99a
also fix the python file. that should be it now, w.r.t. the new token format .bin files
2024-05-21 17:46:14 +00:00
Andrej Karpathy
f671cf9288
more changes, trying to help people out because when this merges to master it will brick everyone's code...
2024-05-20 23:02:41 +00:00
Andrej Karpathy
722e5b2fe5
refactor how we treat datasets, because we're about to have more of them and we don't want them to clutter up root dir etc. this is only step 1, i'm about to refactor a bunch of the dataloading, how the .bin files work and are loaded, how the DataLoader works, etc. This is all needed to support good evals and training at scale
2024-05-20 22:43:40 +00:00
Andrej Karpathy
c2d12f725e
small touchups to grad clip
2024-05-19 17:07:55 +00:00
Erik Schultheis
a9947a8315
added a flag and the missing sqrt;
...
testing now has clipping enabled
2024-05-18 22:47:15 +03:00
Andrej Karpathy
92fc26eba4
the nuts and bolts of gradient accumulation again, merged to master, but there is a bug and it doesn't work, debugging...
2024-05-14 21:32:47 +00:00
Andrej Karpathy
84a2aa0dd0
allow exporting all GPT-2 model sizes now, using --model flag in python script. also allow loading all the models from C, using the new flag -e, to point directly to the file to load. added some error handling for potentially common mistakes. note that loadin GPT-2 XL does not work yet and crashes, but I am imminently merging a fix as the next commit that resolves this
2024-05-08 20:13:14 +00:00
Andrej Karpathy
3cc16f1354
fix logits bug
2024-05-07 00:39:25 +00:00
Horace He
69f6c4f765
Don't return logits during training for PyTorch baseline
...
This improves perf somewhat, since currently it's always returning logits (which thus need to be materialized).
2024-05-06 13:10:17 -07:00
Andrej Karpathy
7d0042d574
make training data configurable and also dont move the engire dataset to device that is not a good idea
2024-05-01 18:15:10 +00:00
Andrej Karpathy
654d6f55c5
add arg to cap the number of steps, and offset all prints to start steps at 1, which i think looks better to the eye
2024-04-30 17:46:36 +00:00
Andrej Karpathy
51face88d9
fix bug where backward/step must be outside of amp context
2024-04-30 15:48:40 +00:00
Andrej Karpathy
f66825faba
allow using flash attention via argparse and print tok/s
2024-04-29 21:05:30 +00:00
Andrej Karpathy
b4c346ac2a
don't hardcode the EOT token and store it in the Tokenizer file
2024-04-29 20:23:41 +00:00
Andrej Karpathy
0e4506ac30
unify these two
2024-04-29 19:03:49 +00:00
ademeure
ebedadfee8
Fix train_gpt2.py to write all weights as BF16 (+merge fixes)
2024-04-29 19:49:57 +01:00
Andrej Karpathy
835060e18d
padded vocab change. touched a lot of code. very stressful and error prone, but i think it is done. had to bump versions on all .bin files, invalidating the previous files. re-run the python training script to re-export the new version files. let's not do too much of things like this in the future lol. actually, fun fact i had a chance to do the padded vocab really really early in the history of llm.c development, and chose not do it, thinking i'll just do it later. i should have done it. such is life, you make mistakes, you accumulate scar tissue, and you learn, and you become better, faster, stronger. this is the mindset one must have to lead a happy and fulfilling life. it's not important that you are perfect at any point in time, it's only important that you keep improving, every day.
2024-04-28 18:47:03 +00:00
Andrej Karpathy
9a91b40207
bring back original ordering. i also had to bump the thresholds by 3X for some tensors and i don't exactly know why sad
2024-04-28 00:02:51 +00:00
Andrej Karpathy
00627070aa
fix a really bad bug in how i was checking the gradients, where i loaded them in the old order, so yeah...
2024-04-27 23:17:22 +00:00
Andrej Karpathy
085270618d
upgrade our .py script so it can also do torchrun with many processes, as our mixed precision mpi nccl code can do now
2024-04-26 19:50:37 +00:00
Andrej Karpathy
8ac4b47696
for now save both fp32 and bf16. todo let's think this through a bit more maybe
2024-04-26 04:36:53 +00:00
Andrej Karpathy
bb56144384
we are an epsilon away from writing our model in bf16 as well, in addition to fp32, with the re-ordered layernorms
2024-04-25 18:32:08 +00:00
Andrej Karpathy
3fb7252924
round 1 of some changes. we will now always write in fp32, even if dtype is set to float16 or bfloat16. next up, we actually want to write in lower precision, when the dtype is set so
2024-04-25 17:38:01 +00:00
Andrej Karpathy
7fe1a4cd49
add device argparse
2024-04-25 16:31:01 +00:00
Andrej Karpathy
98567a4509
turn on benchmarking in mixed precision. example command: python train_gpt2.py --write_tensors 0 --inference_only 0 --dtype bfloat16 --compile 1 --tensorcores 1 --num_iterations 50 --sequence_length 1024
2024-04-24 19:26:40 +00:00