CtrlK
BlogDocsLog inGet started
Tessl Logo

llm-council

Query multiple LLM models in parallel from CodeAct and cross-reference their responses

52

Quality

57%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Passed

No findings from the security scan

Fix and improve this skill with Tessl

tessl review fix ./skills/llm-council/SKILL.md
SKILL.md
Quality
Evals
Security

LLM Council

You can query multiple LLM models with the same prompt directly from CodeAct using the built-in llm_query() and llm_query_batched() functions. Both accept a model= (or models=) keyword that overrides the configured model for that call.

Per-request model override support varies by backend:

BackendHonors model=?Cross-vendor routing?
NEAR AIYesYes (aggregator — hosts models from many vendors)
Anthropic OAuthYesNo (Anthropic models only)
GitHub CopilotYesNo (Copilot-exposed models only)
BedrockNo— (model fixed at construction)
OpenAI / Ollama / Tinfoil via rigNo (silent fallback with warning log)

A genuine cross-vendor council (Anthropic + Google + OpenAI in one batch) therefore only works on an aggregator backend like NEAR AI. On single-vendor backends, use a lineup of models available within that vendor.

When to use a council

  • The user wants diverse perspectives on a question or analysis
  • Cross-referencing answers to increase confidence
  • Comparing reasoning approaches across models
  • Getting a "second opinion" from different AI models
  • Research or evaluation tasks that benefit from multiple viewpoints

Default council line-up

Check the configured backend first (e.g. from LLM_BACKEND or the user's settings) before picking a lineup. Unless the user requests specific models, use the matching default below.

NEAR AI (aggregator — default council):

COUNCIL = [
    "anthropic/claude-opus-4-6",
    "google/gemini-3-pro",
    "zai-org/GLM-latest",
    "openai/gpt-5.4",
]

This 4-model lineup spans the major frontier providers and reasoning styles. It only works on NEAR AI (or another aggregator) — the prefixed model names route inside NEAR AI to the respective vendors.

Anthropic OAuth (Anthropic-only, no cross-vendor routing):

COUNCIL = [
    "claude-opus-4-6",
    "claude-sonnet-4-6",
    "claude-haiku-4-5",
]

Use different Anthropic tiers for diversity of reasoning depth vs. speed.

GitHub Copilot (whatever Copilot exposes):

COUNCIL = ["gpt-5.4", "claude-opus-4-6", "gemini-3-pro"]

Copilot's available models shift over time — call llm_query with the user's configured default if you're unsure which are reachable.

Bedrock / OpenAI / Ollama / Tinfoil: per-request model= is not honored by these backends. A council is not possible without switching backends — tell the user and fall back to a single-model answer.

If the user names specific models, always use those instead of the defaults.

API

Single call with a model override

answer = llm_query(
    prompt="What is X?",
    context="Optional background",     # optional
    model="anthropic/claude-opus-4-6", # optional per-call override
)

Parallel council (same prompt, many models)

COUNCIL = [
    "anthropic/claude-opus-4-6",
    "google/gemini-3-pro",
    "zai-org/GLM-latest",
    "openai/gpt-5.4",
]
responses = llm_query_batched(
    prompts=["What are the main risks of X?"] * len(COUNCIL),
    models=COUNCIL,                    # parallel array, length must match prompts
    context="Answer in 3-5 bullet points.",
)
# `responses` is a list of strings in the same order as `models`.
# If a specific model is unavailable, that slot returns "Error: ..." —
# the rest of the batch still completes.

Single model applied to many prompts

results = llm_query_batched(
    prompts=["Q1", "Q2", "Q3"],
    model="anthropic/claude-opus-4-6", # singular: applies to every prompt
)

Mixing models= slots with None

A None slot inside models=[...] means "no override for this prompt" — that call uses the configured default model. The singular model= kwarg does NOT backfill None slots; it is only used when models= is omitted entirely.

After collecting responses

  1. Identify consensus — note where models agree
  2. Flag disagreements — analyze where they diverge and why
  3. Synthesize — produce a unified answer that accounts for all perspectives
  4. Cite — reference which model contributed each insight

Full example

COUNCIL = [
    "anthropic/claude-opus-4-6",
    "google/gemini-3-pro",
    "zai-org/GLM-latest",
    "openai/gpt-5.4",
]
question = "What are the main risks of relying on a single LLM provider?"

responses = llm_query_batched(
    prompts=[question] * len(COUNCIL),
    models=COUNCIL,
    context="Answer concisely in 3-5 bullet points.",
)

# Build a synthesis prompt
labelled = "\n\n---\n\n".join(
    f"**{model}**:\n{resp}" for model, resp in zip(COUNCIL, responses)
)

synthesis = llm_query(
    prompt=f"Synthesize a balanced answer from these {len(COUNCIL)} expert opinions:\n\n{labelled}",
    context="Identify consensus, flag disagreements, and produce a unified answer.",
)

FINAL(synthesis)

Notes

  • llm_query_batched() runs all calls in parallel — total latency is roughly the slowest model.
  • If models= is provided, it must be the same length as prompts.
  • Use model= (singular) when you want one model applied to every prompt; use models= (plural list) for the council pattern.
  • Errors from individual models (unavailable model, provider timeout, etc.) surface as "Error: ..." strings in the result list, so a single bad model does not fail the whole batch. Argument validation errors — wrong types, or models/prompts length mismatch — still raise exceptions.
Repository
nearai/ironclaw
Last updated
First committed

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.