徽标
我的项目 WeHub 热榜 数据中心
注册 登录
徽标
项目文件夹 工单
mirrors / liguodongiot--llm-action
项目文件夹 工单

项目文件夹

liguodongiot--llm-action
liguodongiot--llm-action mirrors
我的项目 创建
文件
main
/llm-inference/offload.md
T
wehub-resource-sync 605accf02b chore: import upstream snapshot with attribution
2026-07-13 13:23:29 +08:00

638 B
原始文件 永久链接 Blame 文件历史

  • https://huggingface.co/docs/accelerate/concept_guides/big_model_inference

  • https://huggingface.co/docs/transformers/big_models

  • Efficient and Economic Large Language Model Inference with Attention Offloading

  • https://arxiv.org/pdf/2405.01814

  • DeepSpeed Inference: Enabling Efficient Inference of Transformer Models at Unprecedented Scale

  • https://arxiv.org/pdf/2207.00032

FlexFlow

  • https://github.com/flexflow/FlexFlow

FlexGen

  • https://github.com/FMInference/FlexGen

kv cache offload:

  • https://github.com/NVIDIA/TensorRT-LLM/blob/a96cccafcf6365c128f004f779160951f8c0801c/docs/source/kv_cache_reuse.md
在新工单中引用 查看 Git Blame 复制永久链接
由 wehub 强力驱动
简体中文
English 简体中文
许可证 API