最終更新:2025-05-19 (月) 08:31:31 (448d)  

LLM/バックエンド
Top / LLM / バックエンド

llama.cpp

vLLM

Text Generation Inference

ExLlama

  • A standalone Python/C++/CUDA implementation of Llama for use with 4-bit GPTQ weights, designed to be fast and memory-efficient on modern GPUs.

ExLlamaV3?

ExLlamaV2

  • an inference library for running local LLMs on modern consumer GPUs.

Ollama

MLC LLM

TensorRT-LLM

Introducing MII?

関連