AI & ML interests
One-click deployment of Open-source LLMs, on managed and dedicated GPUs.
Recent Activity
Organizations
view article Deploying Qwen3.8-2.4T-A95B with vLLM: Verified GPU Pods, Quants, and Serving Recipes
hexgridcloud
• • 1
published an article about 1 month ago view article Qwen3.6-27B-FP8 on One RTX 6000 Ada: Fast TTFT, 314 tok/s Decode Generation [Benchmark]
published an article about 2 months ago view article Gemma-4 31B + vLLM on RTX 6000 PRO : A Real-Load Benchmark
hexgridcloud
• • 4