Gruve AI
- Product
- Gruve AI, managed inference provider built on distributed spot compute
- Role
- Product, 4-person team working directly with the VP of Product and CEO
- Skills
- Cost modelingPrototypingAI infrastructure economicsCustomer discoveryCompetitive researchB2B product strategy
The Problem
Gruve had raised $50M for a new inference vertical with a specific thesis: GPU capacity sits idle all over the world, spot pricing makes it cheap, and a managed inference provider built on that distributed supply could undercut the incumbents. The vertical has since spun out as its own company, AI Fabrik.
The thesis was solid on supply and untested on demand. Managed inference is crowded. OpenRouter runs a marketplace across every API, Fireworks and Together AI compete on serving performance, and cheap capacity does not tell you who switches provider, when they switch, or what they think they are buying.
The System
Four of us, all Stanford grads, working directly with the VP of Product and the CEO. Weekly cadence: run customer and competitor calls, then turn what we heard into candidate product features that could work as an internal tool and as an acquisition channel for Gruve's GPUs at the same time.
I led the initial market research and customer discovery. We mapped the incumbents, then went to the people who run this for a living, LLM inference engineers and GTM leads, to understand how they price, where margin actually comes from, and what customers complain about. Most of those conversations landed on SLAs, and how stringent a guarantee a buyer insists on before moving production traffic.
That produced the segmentation the product got built around. Early-stage customers fixate on latency and SLA guarantees and barely look at cost. Cost becomes the deciding factor as they scale. Which pointed at three segments where that crossover arrives early and hard: video, voice AI, and agentic companies, where usage is continuous and cost and performance are both first-order concerns.
So we built a beta cost projection platform for voice and video AI. Top-down rather than generic: pick a real customer type and use case, video intelligence for security and voice for customer service, model the actual real-time workload, then run it across generic and specialized models at different quantization levels. The output shows how cost scales exponentially with usage and over time, which is the number a founder cannot get off a price-per-token page.

The Insight
Managed inference gets bought on a latency and SLA guarantee first. Price becomes the argument only once volume makes it one. For a provider entering on a cost advantage, that reorders the pitch: lead with reliability, and make the cost case inside the customer's own workload rather than in list pricing.
It is also why the projection tool worked as a go-to-market asset rather than only an internal model. Showing a voice AI founder what their inference bill looks like at ten times current usage starts the conversation on their problem instead of our capacity.