Wedoo
LLM Manager
No model loaded
Copy Endpoint
Unload
▼
No output yet
External llama-server processes
Configure
Server
Loading...
Port
Context Size (-c)
Predict / Max Tokens (-n)
Threads (-t)
Batch Threads (-tb)
GPU Layers (-ngl)
Main GPU (-mg)
Split Mode (-sm)
layer
row
none
CUDA Devices (-dev)
Tensor Split (-ts)
KV Cache K (-ctk)
KV Cache V (-ctv)
Batch Size (-b)
Micro Batch (-ub)
Parallel (-np)
Reasoning (-rea)
auto
on
off
Reasoning Budget (-1=unrestricted)
Cache RAM (-cram, MiB)
API Key
Use mmproj (auto-detect)
mmproj path
Flash Attention
Cont Batching
mlock
no-mmap
KV Offload
Fused MoE
Fused Up-Gate
Grouped Expert Routing
CPU MoE (all experts on CPU)
Defer Expert mmap
Metrics
MoE layers on CPU (--n-cpu-moe, 0=off)
Load Model
Loading models...