最終更新:2025-05-19 (月) 08:31:31 (448d)
LLM/バックエンド
llama.cpp
vLLM
Text Generation Inference
ExLlama
- A standalone Python/C++/CUDA implementation of Llama for use with 4-bit GPTQ weights, designed to be fast and memory-efficient on modern GPUs.
ExLlamaV3?
ExLlamaV2
- an inference library for running local LLMs on modern consumer GPUs.

