HUM AI FR FriendliAI Tech & Research · FriendliAI Improve Latency and Throughput with Weight-Activation Quantization in FP8 Quantization is a popular technique used to reduce the size of a machine learning model by lowering the numerical precision of some of its…
AI TCH FR FriendliAI Tech & Research · FriendliAI The Essential Checklist: Fix 6 Common Errors When Sharing Models on Hugging Face Hugging Face has become the de facto platform for sharing open-source AI models. But you need to be careful when uploading your models if…
AI FR FriendliAI Tech & Research · FriendliAI How to Use Hugging Face Multi-LoRA Adapters In the previous article, we explored the mechanics behind LoRA (Low-Rank Adaptation) and its growing importance in adapting large-scale…
AI FR FriendliAI Tech & Research · FriendliAI LG AI Research Partners with FriendliAI to Launch EXAONE 4.0 for Fast, Scalable API We’re thrilled to announce our latest strategic partnership with LG AI Research to bring their state-of-the-art language model, EXAONE 4.0…