Reposted by Thanos Stratikopoulos
Can Java compete with llama.cpp/ollama? YES! With @jitllm.io!
jitLLM uses NVIDIA CUDA features and libraries (cuBLAS, cuDNN) and reaches 0.90× to 1.12× of llama.cpp's throughput!
jitLLM is empowered by @tornadovm.org and will be soon in @langchain4j.dev and @quarkus.io.
github.com/beehive-lab/...