LLM speculative inference server for consumer & heterogeneous hardware
-
Updated
Aug 15, 2026 - C++
LLM speculative inference server for consumer & heterogeneous hardware
BestBuy Bot is an Add to cart and Auto Checkout Bot. This auto buying bot can search the item repeatedly on the ITEM page using one keyword. Once the desired item is available it can add to cart and checkout very fast. This auto purchasing BestBuy Bot can work on Firefox Browser so it can run in all Operating Systems. It can run for multiple ite…
Pre-built wheels for llama-cpp-python across platforms and CUDA versions
Best Buy Bullet Bot, abbreviated to 3B Bot, is a stock checking bot with auto-checkout created to instantly purchase out-of-stock items on Best Buy once restocked. It was designed for speed with ultra-fast auto-checkout, as well as the ability to utilize all cores of your CPU with multiprocessing for optimal performance.
Open source DIY AI computing platform: Build a powerful RTX 3090 GPU rig under €1,300 for local LLM inference, training, and AI development. Complete with parts list, assembly guide, Ubuntu setup, and remote access configuration. Ideal for students, researchers, and hobbyists wanting affordable AI hardware without cloud dependencies.
🌐 Automate and manage BitBrowser tasks for efficient Google One student discount processing in one streamlined system.
Full benchmark traces: Laguna S 2.1 INT4 vs DFlash speculative decoding on 4x RTX 3090 (vLLM 0.25, TP4, 96GB VRAM). 152 measurements, raw JSON + GPU telemetry.
PlantCare AI | High-performance Plant Disease Detection system optimized for RTX 3090. Identifies 38 diseases with 93% accuracy. Features a futuristic Glassmorphism UI, real-time diagnosis, and AI-driven treatment recommendations. 🌿🌱
poolside Laguna-S-2.1 INT4 + DFlash speculative decoding on 4x RTX 3090: 200K context, 282 tok/s peak decode, gate-proven with a 190K-token prompt. Full levers menu + failure catalog.
Dynamic open-source Google Sheets tables inspired by the famous Hive Systems ones. Just input the H/s/GPU data.
Benchmark llama.cpp prefill/decode speed at real context depth, not on an empty context. Includes a Qwen3.6-27B run on 1x and 2x RTX 3090.
Add a description, image, and links to the rtx3090 topic page so that developers can more easily learn about it.
To associate your repository with the rtx3090 topic, visit your repo's landing page and select "manage topics."