Tag
Language Models.
Latest
A Smoke Test of ZML's llmd: DFlash Speculative Decoding on Consumer AMD GPUs.
A first-contact test of ZML's alpha inference server on RDNA4 hardware their docs don't mention testing against: it works, and DFlash speculative decoding on Gemma-4 gave a real 1.66x decode speedup.
Read the article →
All writing
OpenLLMetry vs. OpenInference: The Observability Stack Your LLM App Needs
Your app looks fine on the surface, but users say it's failing. Traditional logs can't explain why. Enter OpenLLMetry and OpenInference—tools that bring clarity to AI observability by capturing rich, standardized telemetry from your LLM stack.
The Unseen Revolution: Why Small Models could be the True Engine of Agentic AI
Agentic AI is transforming the field, enabling fast, efficient, and affordable automation far beyond chatbots through small, specialized models. These Small Language Models (SLMs) can outperform large models in real-world tasks, powering the next generation of enterprise and on-device AI.
No posts match "".


