CtrlK
BlogDocsLog inGet started
Tessl Logo

fireworks-ai-inference

Fast inference and fine-tuning platform with serverless and on-demand GPU deployments. OpenAI-compatible API for chat completions, embeddings, function calling, vision, and structured output. Supports SFT, DPO, and RL fine-tuning. SOC2 + HIPAA compliant.

54

Quality

61%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Passed

No findings from the security scan

Fix and improve this skill with Tessl

tessl review fix ./backend/cli/skills/cloud-compute/fireworks-ai/SKILL.md
SKILL.md
Quality
Evals
Security

Quality

Content

57%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

A dense, largely executable reference with excellent troubleshooting and pricing tables, but it is a monolithic 670-line document that should offload API, CLI, and catalog detail to reference files, repeats its client-setup pattern multiple times, and lacks validation steps around destructive deployment and batch operations. Actionability is its strongest dimension; progressive disclosure is structurally limited by having no bundle at all.

Suggestions

Split the model catalog, pricing tables, CLI reference, and fine-tuning API details into references/ files (e.g. MODELS.md, PRICING.md, FIRECTL.md, FINE-TUNING.md), keeping only quick start and key examples in SKILL.md.

Define the OpenAI client once and drop the duplicate chat/streaming/embeddings examples in the 'OpenAI Compatibility' section, replacing them with a one-line note that only base_url and api_key change.

Add validation checkpoints around destructive and batch operations — e.g. check deployment state before delete, verify batch file format before upload, and confirm job state before stopping a fine-tuning job.

DimensionReasoningScore

Conciseness

The body is mostly dense tables and executable code with little conceptual padding, but the OpenAI client setup block repeats four times and the 'OpenAI Compatibility' section re-demonstrates chat, streaming, and embeddings already covered in 'Inference' and 'Embeddings'; minor concept explanations (DPO, reward models) add little. Anchor 3 — mostly efficient but could be tightened — rather than 2 (no heavily padded sections) or 4 (the duplicated setup and examples exceed 'minor').

3 / 5

Actionability

Nearly all snippets are executable and specific (streaming loop, tool schema, response_format with JSON schema, firectl commands, troubleshooting pairs), but the requests-based fine-tuning/deployment snippets depend on an undefined `headers` variable and contain `{account_id}`/`{dataset_id}` placeholders, so they are not copy-paste ready without assembly. Anchor 4 — mostly executable with minor gaps — rather than 5, which demands fully copy-paste-ready coverage of common cases.

4 / 5

Workflow Clarity

Sequences exist (create job → monitor job; create deployment → scale/delete) and the Common Issues table offers some error recovery, but destructive operations (deployment delete, fine-tuning-job stop) and batch operations carry no validation or verification steps. Per the rubric's explicit cap, missing validation in destructive/batch workflows caps workflow clarity at 3.

3 / 5

Progressive Disclosure

No bundle files exist (references/, scripts/, assets/ are absent), so the full ~670-line document — API reference, pricing tables, model catalog, CLI docs — is inlined in SKILL.md. Section headers are well organized, but bulk reference material that clearly belongs in separate files is inline, matching anchor 3's example of 200+ lines of API reference inlined. Not 2, since structure is present rather than 'minimal'; not 4, since nothing is split out.

3 / 5

Total

13

/

20

Passed

Description

66%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

A specific, capability-dense description that clearly states what the platform does, but it omits any 'use when' trigger guidance and never names the platform itself, weakening both completeness and distinctiveness against sibling inference-provider skills. Trigger vocabulary is good though not exhaustive.

Suggestions

Append an explicit trigger clause, e.g. 'Use when the user mentions Fireworks AI, needs fast serverless or dedicated-GPU inference for open models, or wants SFT/DPO/RL fine-tuning without infrastructure.'

Include natural synonym triggers users actually say — 'LLM inference', 'open models', 'LLM API', 'GPU hosting' — to lift trigger-term coverage.

Name the platform ('Fireworks AI') in the description to reduce conflict risk with other OpenAI-compatible providers like Together or Groq.

DimensionReasoningScore

Specificity

Lists several concrete capabilities ('chat completions, embeddings, function calling, vision, and structured output', 'SFT, DPO, and RL fine-tuning', 'serverless and on-demand GPU deployments') with only minor gaps (batch API, prompt caching absent; 'Fast' is a marketing lead-in, not an action). Matches anchor 4 — lists several specific actions with minor coverage gaps — but not anchor 5, which demands comprehensive concrete-action coverage free of that kind of padding.

4 / 5

Completeness

The 'what' is explicit and concrete, but there is no 'Use when...' clause or equivalent trigger guidance anywhere, so per the rubric guideline a missing 'Use when' caps completeness at 3 ('what' clear, 'when' missing or only weakly implied). Score 4 would require an explicit, if imperfect, 'when' clause.

3 / 5

Trigger Term Quality

'inference', 'fine-tuning', 'chat completions', 'embeddings', 'function calling', 'structured output', and 'GPU deployments' are terms users naturally say when needing this skill. It falls short of anchor 5's comprehensive synonym coverage — no 'LLM', 'language models', 'open models', or brand-name trigger — but is well above anchor 3's 'missing common variations'.

4 / 5

Distinctiveness Conflict Risk

'SOC2 + HIPAA compliant', 'on-demand GPU deployments', and 'SFT, DPO, and RL fine-tuning' carve a fairly distinct niche, but 'OpenAI-compatible API for chat completions, embeddings, function calling, vision' overlaps with every other OpenAI-compatible inference provider (Together, Groq, vLLM). Mostly distinct with minor overlap risk against closely related skills — anchor 4, not 5, since the brand name itself is absent from the description.

4 / 5

Total

15

/

20

Passed

Validation

87%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 14 / 16 Passed

Validation for skill structure

CriteriaDescriptionResult

skill_md_line_count

SKILL.md is long (679 lines); consider splitting into references/ and linking

Warning

frontmatter_unknown_keys

Unknown frontmatter key(s) found; consider removing or moving to metadata

Warning

Total

14

/

16

Passed

Repository
synthetic-sciences/openscience
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.