README: The main goal of `llama.cpp` is to enable LLM inference with minimal setup and state-of-the-art performance on a wide range of hardware - locally and in the cloud.
README: Multimodal support arrived in `llama-server`: #12898 | documentation
README: Run with Docker - see our Docker documentation