https://github.com/vllm-project/llm-compressor/
https://nvidia.github.io/TensorRT-Model-Optimizer/guides/_choosing_quant_methods.html
https://github.com/NVIDIA/TensorRT-Model-Optimizer/blob/main/examples/deepseek/README.md