文件历史

提交图

253 次代码提交

作者 SHA1 备注 提交日期
Erik Schultheis d6793645f5 replace cublaslt with custom matmul 2024-06-12 09:42:19 +03:00
Erik Schultheis 7f3f3acf16 shared memory matmul 2024-06-11 21:03:43 +03:00
Erik Schultheis 071c977c79 vectorized memory access 2024-06-11 00:46:51 +03:00
Erik Schultheis bb89efdcb1 register reuse 2024-06-11 00:41:08 +03:00
Andrej 637c1b69c4 Merge pull request #552 from karpathy/feature/streams
Feature/streams
2024-06-08 09:25:56 -07:00
Andrej Karpathy ee6b3c9325 take out Async copies and memsets 2024-06-08 16:17:02 +00:00
Aleksa Gordic 7b99ac96a5 Fix mem leaks, reduce memory, refactor 2024-06-06 19:58:50 +02:00
Andrej d931fc0507 -std=c++17 default, merge PR #559 from ngc92/ubuntu20.04
set compile flag and add ci check
2024-06-06 09:23:07 -07:00
Erik Schultheis 875196b672 add c++17 flag also to dev/cuda 2024-06-06 19:02:45 +03:00
Aleksa Gordic c116fbfa29 Fix q,k bug, move all kernels to q@k^t form 2024-06-05 22:22:55 +02:00
Aleksa Gordic 875d362381 Refactor trimat 2024-06-05 17:33:26 +02:00
Andrej 80e0c3b8d9 Merge pull request #512 from gordicaleksa/refactor_encoder_bwd_kernel
Remove redundant CPU computation in encoder bwd
2024-06-02 16:43:10 -07:00
Aleksa Gordic 290c00a362 Remove redundant CPU computation 2024-06-01 18:27:24 +02:00
Andrej a26041af12 Merge pull request #495 from ChrisDryden/shared_memory
Removed unnecesary shared memory due to blockreduce using static defined shared memory
2024-06-01 08:12:21 -07:00
Aleksa Gordic 8eafd40615 Update compile cmd in the dev/cuda README 2024-06-01 16:46:32 +02:00
Christopher 5450632237 Removed unnecesary shared memory due to blockreduce using static defined shared memory 2024-05-30 03:01:52 +00:00
Andrej 55f3665b64 Merge pull request #493 from vyom1611/master-2
Modal benchmarking script updated to replace deprecated calls
2024-05-29 07:50:10 -07:00
vyom1611 062f096da1 Modal benchmarking script updated to replace deprecated calls 2024-05-29 20:01:41 +05:30
lancerts 47c670b885 amend 2024-05-27 15:13:11 -07:00
lancerts c4c985cb87 spell explicitly uint to unsigned int 2024-05-27 14:44:47 -07:00
lancerts f27ca4df41 fix the issue Mismatch of dweight at layernorm_backward.cu 2024-05-27 14:39:57 -07:00
Erik Schultheis 7dc3b7d7dc bugfix 2024-05-27 16:36:44 +03:00
Erik Schultheis 5f73ecfcf1 fix C % 256 != 0 2024-05-27 13:15:36 +03:00
Erik Schultheis d35daf19c9 fail fast and hard; don't go into the deadlock 2024-05-27 13:15:36 +03:00
Erik Schultheis 3b6808269a some comments and optimized shared memory amount 2024-05-27 13:15:36 +03:00
Erik Schultheis 9eba86f0a5 fully vectorized smem access 2024-05-27 13:15:36 +03:00
Erik Schultheis cb8cc25e3a utilities for Packed and more vectorization 2024-05-27 13:15:36 +03:00
Erik Schultheis 53ee3297dd use vectorized access to shared memory 2024-05-27 13:15:36 +03:00
Andrej 6a7fd56d7c Merge pull request #439 from lancerts/matmul-fix
Fix the unsupported block_size in matmul_backward_bias kernel 1
2024-05-24 17:26:37 -07:00
Andrej Karpathy dbacaf84cf Merge branch 'deterministic_layernorm' of https://github.com/ademeure/llm.c into ademeure-deterministic_layernorm 2024-05-24 21:46:25 +00:00
Erik Schultheis 1b98637960 int -> int64_t 2024-05-24 20:11:34 +03:00
ademeure 7cbeefc7f3 added new layernorm backward to /dev/cuda/ 2024-05-21 23:26:54 +01:00
lancer 2b0667aee1 update the utils function and assert 2024-05-20 08:00:39 -07:00
lancer 6348d4196d fix the unsupported block_size 2024-05-19 17:39:25 -07:00
Andrej Karpathy c2d12f725e small touchups to grad clip 2024-05-19 17:07:55 +00:00
Erik Schultheis 66ce5766e0 fixed up dev/cuda 2024-05-18 23:06:44 +03:00
Erik Schultheis d7a81ef26f added a useful mixed precision utility for dev/cuda 2024-05-18 22:45:44 +03:00
Andrej 109f516367 Merge pull request #383 from KarhouTam/feature/online-softmax-forward-without-cgs
Free cooperative groups implementation of online softmax forward
2024-05-16 21:23:56 +01:00
Erik Schultheis 858c6e6dae deterministic kernel 2024-05-16 01:04:10 +03:00
Erik Schultheis 2ccdfb70e0 general cleanup 2024-05-16 01:04:10 +03:00
Andrej Karpathy 2d43e5bc97 remove legacy comment 2024-05-14 19:14:07 +00:00
Andrej 222d59fa2f Merge pull request #408 from ngc92/layernorm-bw-dev-cuda
Layernorm backward updates
2024-05-14 20:13:09 +01:00
Erik Schultheis e553e2f084 update dev/cuda/layernorm_backward and improve validate_result to take into account fp epsilon when comparing results 2024-05-14 20:44:50 +03:00
Erik Schultheis dd8c9f5ec9 fix layernorm backward: accumulate weight gradient 2024-05-14 20:43:51 +03:00
Erik Schultheis c66e48c06c fixup comment 2024-05-13 20:58:07 +03:00
Erik Schultheis 65727d5a4d fix CI compile by disabling kernel 5 2024-05-13 19:20:03 +03:00
Erik Schultheis 49ee3c8307 fix non-atomic version:
* accumulate instead of assign
* need dedicated argument to correctly handle the floatX == float case
2024-05-13 18:27:56 +03:00
Erik Schultheis 081d224b21 automatically switch to buffer-less version if that can fill up the GPU 2024-05-13 17:39:32 +03:00
Erik Schultheis c0329ebdba new kernel version with fewer atomics 2024-05-13 17:18:27 +03:00
Erik Schultheis 2287da0120 enable bf16 2024-05-12 19:42:32 +03:00