| | The efficient frontier of LLM inference (baseten.co) |
| 155 points by philipkiely 11 days ago | past | 46 comments |
|
| | Inference Engineering: A free book on the systems behind AI inference (baseten.co) |
| 3 points by DarenWatson 24 days ago | past | 1 comment |
|
| | Inference Engineering by Philip Kiely – Digital Download (baseten.co) |
| 2 points by ajhai 25 days ago | past |
|
| | We built a day-0 API for Kimi K3 (baseten.co) |
| 2 points by philipkiely 47 days ago | past |
|
| | We built the new fastest API for GLM-5.2 (baseten.co) |
| 2 points by philipkiely 49 days ago | past |
|
| | How we built the fastest API for GLM-5.2 (baseten.co) |
| 4 points by Philpax 80 days ago | past |
|
| | How to get GLM 5.2 to 280 tokens per second (baseten.co) |
| 3 points by mikejulietbravo 81 days ago | past | 1 comment |
|
| | We built the fastest API for GLM-5.2 (280 TPS) (baseten.co) |
| 6 points by philipkiely 82 days ago | past |
|
| | Baseten raised a $1.5B Series F and achieved a $13B valuation (baseten.co) |
| 5 points by kodablah 82 days ago | past |
|
| | The Math Behind TurboQuant (baseten.co) |
| 8 points by philipkiely 5 months ago | past | 3 comments |
|
| | Inferless Joins Baseten (baseten.co) |
| 1 point by agcat 6 months ago | past |
|
| | How We Built the Fastest Kimi K2.5 on Artificial Analysis (baseten.co) |
| 3 points by philipkiely 7 months ago | past |
|
| | Continual learning and the post monolith AI era (baseten.co) |
| 1 point by jxmorris12 7 months ago | past |
|
| | Running GPT-OSS-120B at 500 tokens per second on Nvidia GPUs (baseten.co) |
| 247 points by philipkiely on Aug 7, 2025 | past | 175 comments |
|
| | Continuous vs. dynamic batching for AI inference (baseten.co) |
| 1 point by aaronng91 on Aug 6, 2025 | past |
|
| | A guide to LLM inference and performance (baseten.co) |
| 1 point by skidrow on Feb 16, 2025 | past |
|
| | Deploying custom ComfyUI workflows as APIs (baseten.co) |
| 1 point by AnhTho_FR on Nov 20, 2024 | past |
|
| | How to build function calling and JSON mode for open-source and fine-tuned LLMs (baseten.co) |
| 1 point by philipkiely on Sept 12, 2024 | past |
|
| | How to double tokens per second for Llama 3 with Medusa (baseten.co) |
| 2 points by philipkiely on Aug 20, 2024 | past |
|
| | Show HN: Automatically Build Nvidia TRT-LLM Engines (baseten.co) |
| 2 points by mikejulietbravo on Aug 1, 2024 | past |
|
| | Show HN: 60% higher tokens per second for 70B custom LLMs (baseten.co) |
| 1 point by mikejulietbravo on July 31, 2024 | past |
|
| | Show HN: Baseten Chains – Framework and SDK for Multi-Model AI Products (baseten.co) |
| 9 points by mikejulietbravo on June 27, 2024 | past | 5 comments |
|
| | Open Source Inference Engine Baseten Raises $40M from IVP, Spark and Greylock (baseten.co) |
| 2 points by mikejulietbravo on March 14, 2024 | past | 1 comment |
|
| | FP8: Efficient model inference with 8-bit floating point numbers (baseten.co) |
| 2 points by philipkiely on March 8, 2024 | past |
|
| | Introduction to quantizing machine learning models (baseten.co) |
| 1 point by tuhins on Feb 16, 2024 | past |
|
| | Faster Mixtral inference with TensorRT-LLM and quantization (baseten.co) |
| 2 points by tikkun on Dec 27, 2023 | past | 1 comment |
|
| | A guide to open-source LLM inference and performance (baseten.co) |
| 113 points by varunshenoy on Nov 20, 2023 | past | 14 comments |
|
| | How we got Stable Diffusion XL inference to under 2 seconds (baseten.co) |
| 51 points by varunshenoy on Aug 31, 2023 | past | 5 comments |
|
| | SDXL inference in under 2 seconds (baseten.co) |
| 3 points by tuhins on Aug 31, 2023 | past | 1 comment |
|
| | Three techniques to adapt LLMs for any use case (baseten.co) |
| 1 point by philipkiely on June 15, 2023 | past |
|
|
| More |