TECH·4 days agoDeploying a 1-Bit Bonsai-27B Model with PrismML llama.cpp and OpenAI-Compatible Local Inference WorkflowsMarkTechPost
TECH·July 20, 2026Someone Fine-Tuned OpenBMB’s MiniCPM5-1B on Claude Fable 5 Traces to Ship a 657MB Local Thinking ModelMarkTechPost
TECH·May 30, 2026Show HN: Tiny-vLLM – high performance LLM inference engine in C++ and CUDA | Hacker NewsHacker News