CtrlK
BlogDocsLog inGet started
Tessl Logo

groq-cost-tuning

Optimize Groq costs through model routing, token management, and usage monitoring. Use when analyzing Groq billing, reducing API costs, or implementing usage monitoring and budget alerts. Trigger with phrases like "groq cost", "groq billing", "reduce groq costs", "groq pricing", "groq expensive", "groq budget".

78

Quality

100%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Passed

No findings from the security scan

SKILL.md
Quality
Evals
Security

Groq Cost Tuning

Overview

Optimize Groq inference costs through smart model routing, token minimization, and caching. Groq pricing is already extremely competitive, but at high volume the savings from routing classification to 8B vs 70B are 12x per request.

Prerequisites

  • A Groq account with an API key exported as the GROQ_API_KEY environment variable — the groq-sdk client reads it automatically (new Groq()).
  • Node.js with the groq-sdk package installed (npm install groq-sdk).
  • Access to the Groq Console to set spending caps and read the usage dashboard.

Groq Pricing (per million tokens)

ModelInputOutput
llama-3.1-8b-instant~$0.05~$0.08
llama-3.3-70b-versatile~$0.59~$0.79
llama-3.3-70b-specdec~$0.59~$0.99
meta-llama/llama-4-scout-17b-16e-instruct~$0.11~$0.34
whisper-large-v3-turbo~$0.04/hr

Check current pricing at groq.com/pricing.

Instructions

Apply these six levers in order. Each compounds on the last — routing alone is the biggest win (~12x), and caching plus batching halve the remainder. The lean skeleton below shows the routing core; the full code for every step lives in references/implementation.md.

  1. Smart model routing — map each use case to the cheapest model that meets its quality bar (classification/extraction/summarization → llama-3.1-8b-instant; reasoning/code review/chat → llama-3.3-70b-versatile; vision → llama-4-scout).
  2. Minimize tokens per request — trim verbose system prompts and cap max_tokens so a one-word answer never bills for a paragraph.
  3. Batch to reduce overhead — fold many items into one request; 10-in-1 cuts per-request overhead and RPM pressure ~90%.
  4. Cache deterministic requests — at temperature: 0, hash identical prompts into a cache for zero-cost, zero-latency repeat hits.
  5. Usage tracking — log token counts and estimated cost per call to catch spend regressions before the invoice.
  6. Spending limits in console — set a monthly cap, alerts at 50%/80%, and auto-pause in Groq Console > Billing.
import Groq from "groq-sdk";
const groq = new Groq(); // reads GROQ_API_KEY

const ROUTING = {
  classification: "llama-3.1-8b-instant",   // ~$0.05/M
  reasoning:      "llama-3.3-70b-versatile", // ~$0.59/M
};
const getModel = (useCase: string) =>
  ROUTING[useCase] || "llama-3.1-8b-instant";
// Classification on 8B vs 70B = 12x savings

See references/implementation.md for the complete routing table, token-minimization, batching, caching, usage-tracking, and console-limit code.

Output

Applying the workflow produces:

  • A routing map (getModel(useCase)) that resolves every call to the cheapest fit model.
  • A usage log of UsageRecord rows (timestamp, model, prompt/completion tokens, estimated cost) accumulated per call.
  • A daily cost report from dailyCostReport() returning { totalCost, byModel }, e.g. { totalCost: "$2.0000", byModel: { "llama-3.1-8b-instant": "$2.0000" } }.
  • Console spending controls: a monthly cap, 50%/80% alerts, and auto-pause.

Examples

Batch three items in a single call using the batchClassify helper from references/implementation.md:

const labels = await batchClassify([
  "Loved it, five stars",
  "Broke on day one",
  "It was fine, nothing special",
]);
// -> ["positive", "negative", "neutral"]  (1 API call instead of 3)

For the full 100,000-message cost walkthrough and a stacked routing + caching + tracking pipeline, see references/examples.md.

Error Handling

IssueCauseSolution
Costs higher than expected70B for simple tasksRoute classification/extraction to 8B
Spending cap hitBudget exhaustedIncrease cap or reduce volume
Cache not effectiveUnique promptsNormalize prompts before caching
Rate limits causing retriesRPM cap hitBatch requests, spread across time

Resources

Repository
jeremylongshore/claude-code-plugins-plus-skills
Last updated
First committed

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.