Visit websitearrow_forward

Ferrum

Run local LLMs with one Rust binary

Ferrum is an open-source, Rust-native runtime designed for efficient local language model execution on Apple Silicon Metal and NVIDIA CUDA architectures. This single binary offers a complete environment, including an interactive command-line interface and robust APIs compatible with OpenAI's Chat Completions and Responses. Users can operate without dependencies like Python, PyTorch, or vLLM. Key features include direct model-to-API paths, ensuring a streamlined workflow. It supports two distinct accelerator backends: Apple Silicon utilizes Metal with GGUF models, while NVIDIA systems (sm89) leverage CUDA with GPTQ or safetensors models. This flexibility allows a wide range of hardware to efficiently run advanced models locally. Ferrum ensures explicit model selection, giving users full control over which model is active. It also offers comprehensive server controls such as continuous batching, paged KV cache, prefix cache, session cache, and typed admission controls. The platform is rigorously validated through documented release flows, ensuring reliability and stability. Ideal for developers and organizations prioritizing privacy and performance, Ferrum simplifies the process of integrating powerful language models into local applications. Whether for interactive exploration via CLI or serving models over HTTP, it provides a performant and lightweight solution.
local_fire_department
Find trending agents & tools
star_shine
Compare options without overload
database
Over 20000 results
local_fire_department
Find trending agents & tools
star_shine
Compare options without overload
database
Over 20000 results
local_fire_department
Find trending agents & tools
star_shine
Compare options without overload
database
Over 20000 results
local_fire_department
Find trending agents & tools
star_shine
Compare options without overload
database
Over 20000 results
share
Rate and share your findings
refresh
Refine and run another iteration
check
Only 4 focused results per step
share
Rate and share your findings
refresh
Refine and run another iteration
check
Only 4 focused results per step
share
Rate and share your findings
refresh
Refine and run another iteration
check
Only 4 focused results per step
share
Rate and share your findings
refresh
Refine and run another iteration
check
Only 4 focused results per step

Search AI solutions for your tasks

Artificial intelligence agents & tools automate your business processes in +1000 knowledge domains
Find productsstar_shine