simplellmfunc
Use SimpleLLMFunc to build typed async LLM functions, chat agents, tools, event-stream consumers, and PyRepl/selfref workflows. Use when writing or editing app code that imports SimpleLLMFunc, configures provider.json, adds @llm_function/@llm_chat/@tool, chooses OpenAICompatible or OpenAIResponsesCompatible, or mounts PyRepl/FileToolset.
SimpleLLMFunc Usage
When to use this skill
- Use this skill for application-level work built on top of SimpleLLMFunc.
- Use it when the task mentions
llm_function,llm_chat,tool,OpenAICompatible,OpenAIResponsesCompatible,provider.json,reasoning,enable_event,PyRepl,SelfReference,FileToolset, or the built-in TUI. - Do not use this skill for framework-internal refactors; use
simplellmfunc-developerfor that.
Core philosophy
- Treat the LLM call like a normal Python function call.
- Put the prompt in the function docstring.
- Let Python parameters and return annotations define the contract.
- Keep orchestration in Python instead of hiding it in giant prompt strings.
- Prefer small, typed, composable building blocks.
How system prompts are really built
This is critical for writing good prompts in SimpleLLMFunc: your docstring is important, but it is usually not the whole final system prompt.
llm_function
- Your docstring is first treated as
function_description. - If you pass
_template_params, the docstring is formatted before prompt construction. - The framework then wraps that docstring inside a system template that also adds:
- parameter type descriptions
- return-type instructions
- plain-text or XML output constraints depending on the return type
- If tools are mounted, the framework prepends a deduplicated
<tool_best_practices>block before the main system prompt. - Runtime argument values are not put in the system prompt; they go into the user prompt.
Write llm_function docstrings as task policy, quality bar, constraints, and style guidance. Do not waste docstring space restating parameter schemas or low-level output formatting that the framework already injects.
llm_chat
- Base system prompt is either:
- the latest
systemmessage inhistory, if one exists, or - the function docstring otherwise
- the latest
- Then the framework prepends
<tool_best_practices>when tools exist. - Then it appends a
<must_principles>block that tells the model to use native structured tool calls instead of writing fake tool calls in assistant text. - Older
systemmessages from history are filtered out; only the latest one wins. - Current turn data is added as the user message, not merged into the system prompt.
Write llm_chat docstrings as stable assistant policy and long-lived behavior. Put current task content in the function call arguments, not in the docstring.
Prompt-writing implications
- For
llm_function, think: function contract + execution strategy. - For
llm_chat, think: assistant identity + durable rules. - Put tool-usage advice in tool
best_practiceswhen possible, not only in the main docstring. - If you need to durably change a chat agent's context mid-run, use the latest history
systemmessage or self-reference context helpers such asruntime.selfref.context.remember(...)/runtime.selfref.context.compact(...)instead of trying to mutate old docstrings. - When
llm_chatis bound toSelfReference, the framework syncs turn state automatically through ReAct lifecycle hooks. Treatruntime.selfref.context.*as the supported way to change durable context instead of editing old history messages in place.
Fast start modes
- Project mode: load models from
provider.jsonwithOpenAICompatible.load_from_json_file(...)orOpenAIResponsesCompatible.load_from_json_file(...)when you have shared config or multiple models. - Instant mode: write a tiny
python - <<'PY'snippet, constructAPIKeyPoolplus the right adapter (OpenAICompatibleorOpenAIResponsesCompatible) directly, decorate one function, and call it immediately. - Prefer instant mode for quick shell usage, generated scripts, demos, and one-off local agents.
Interface choice
- Use
OpenAICompatiblefor normal OpenAI-style chat/completions endpoints. - Use
OpenAIResponsesCompatiblefor OpenAI Responses API endpoints when you want the Responses transport while keeping the same decorator surface. OpenAIResponsesCompatibleuses the sameprovider.jsonshape and direct-construction shape asOpenAICompatible.- When you use
OpenAIResponsesCompatible, the framework still builds normal chat/system messages first; the adapter maps the chosen system prompt to Responsesinstructionsand forwardsreasoning={...}kwargs. - Keep prompt authoring the same across both adapters. Do not rewrite docstrings around raw Responses wire format.
Export the packaged skill
After installing SimpleLLMFunc, export the bundled Agent Skills with:
simplellmfunc-skill usage ~/.config/opencode/skills
simplellmfunc-skill developer ~/.config/opencode/skills
usageexports thesimplellmfuncfolder.developerexports thesimplellmfunc-developerfolder.- The second argument is the parent directory that receives the exported skill folder.
- Add
--forceif you need to overwrite an existing exported copy.
Configuration essentials
Minimal provider.json shape
SimpleLLMFunc expects provider.json to be:
{
"openrouter": [
{
"model_name": "z-ai/glm-5",
"api_keys": ["sk-key-1", "sk-key-2"],
"base_url": "https://openrouter.ai/api/v1",
"max_retries": 5,
"retry_delay": 1.0,
"rate_limit_capacity": 20,
"rate_limit_refill_rate": 3.0
}
]
}
- Top level = provider id -> model config list.
- Lookup shape after loading =
providers[provider_id][model_name]. - Start from
examples/provider_template.jsonwhen possible.
How to organize provider.json
Treat provider.json as the canonical project-level model routing table, not as a random dump of keys.
Recommended organization rules:
- Group by provider first, then keep a short list of model configs under that provider.
- Keep
model_namevalues stable and unique within one provider. - Put multiple keys under the same hot model in
api_keysinstead of duplicating the model entry. - Tune
max_retries,retry_delay,rate_limit_capacity, andrate_limit_refill_rateper model, not once globally. - Keep the file focused on runtime model access concerns only: provider, model, keys, endpoint, retry, and rate limits.
- Use
OpenAICompatible.load_from_json_file(...)once near application setup, then pass resolved model handles into decorators.
Recommended mental model:
provider.jsondecides what model surface is available.- typed decorators and tools decide how tasks are expressed.
- your harness decides what context reaches the model for one task.
.env and environment variables
The framework mainly reads .env / environment variables for logging and optional Langfuse observability.
LOG_LEVEL=WARNING
LOG_DIR=logs
LANGFUSE_PUBLIC_KEY=your_public_key
LANGFUSE_SECRET_KEY=your_secret_key
LANGFUSE_BASE_URL=https://cloud.langfuse.com
LANGFUSE_EXPORT_ALL_SPANS=true
LANGFUSE_ENABLED=true
- Precedence is: runtime environment variables ->
.env-> framework defaults. - Recommended default:
LOG_LEVEL=WARNINGto reduce noisy framework logs during normal agent usage. provider.jsonis the main model/provider config file;.envis not the primary place for provider definitions in project mode.- For shell-first one-offs, direct constructor snippets are still fine even without
provider.json.
Default workflow
- Choose the right surface:
@llm_functionfor one typed call.@llm_chatfor multi-turn agent behavior.@toolfor external capabilities the model may call.PyReplwhen the model needs persistent code execution or runtime primitives.
- Define everything as
async def. - Write a precise docstring prompt and leave the body as
pass. - Use explicit parameter types and a typed return value; prefer Pydantic for structured outputs.
- Build the model either from
provider.jsonor directly withAPIKeyPool+OpenAICompatible/OpenAIResponsesCompatible. - Add
toolkit=[...]only when the task truly needs tools. - Validate with a focused example or event consumer.
Strong typing and Pydantic (Recommended default)
Prefer explicit, typed function contracts over loose string-in/string-out prompting.
Recommended rules:
- Use narrow, explicit parameter types.
- Use typed return values whenever the output shape matters.
- Prefer Pydantic models for structured outputs that must be stable, inspectable, or reusable.
- Let the framework derive structured parsing from the Python return type instead of asking the model to hand-roll JSON.
- Keep the docstring focused on task intent, quality bar, constraints, and edge cases, not schema duplication.
Good pattern:
from pydantic import BaseModel, Field
from SimpleLLMFunc import llm_function
class SearchSummary(BaseModel):
answer: str = Field(description="Direct answer to the user's question")
evidence: list[str] = Field(description="Supporting evidence bullets")
confidence: float = Field(description="0.0 to 1.0 confidence score")
@llm_function(llm_interface=llm)
async def summarize_search_results(query: str, snippets: list[str]) -> SearchSummary:
"""
Answer the query using the provided snippets.
Prefer concise claims backed by evidence.
If the snippets are insufficient, lower confidence and say so explicitly.
"""
pass
Prefer this over:
- returning a free-form string and parsing it manually later
- asking the model to invent ad hoc JSON output shapes in the docstring
- mixing many unrelated fields into one loose
dict[str, Any]unless the shape is genuinely variable
Use plain str returns only when free-form text is truly the product.
Harness Engineering: context planning comes first
When building on SimpleLLMFunc, think in terms of harness engineering.
Core philosophy:
- the main job is context planning
- an agent is not a person; it is a method for constructing the right context for each reasoning step
- at every step, the model should see the shortest clean context that is still complete for the current task
- the system should be designed so this context can be maintained in a durable closed loop
What this means:
- when an agent fails, first ask what was missing from the environment or context
- encode the fix into the environment itself: tools, checks, constraints, docs, files, or workflow structure
- do not rely on the operator remembering the lesson in their head
Core beliefs:
- model capability is usually not the bottleneck; context quality is
- broad context stuffing is usually worse than precise context planning
- agent failures are predictable and should map to concrete harness changes
- cross-session memory must be rebuilt from external state, not assumed
Practical implications:
- choose
@llm_functionwhen one typed transformation gives the cleanest context - choose
@llm_chatonly when multi-turn state genuinely improves the context for the task - decide between keeping history, forking, or using a sub-agent based only on which option yields the most accurate and compact context for the current reasoning step
- persist state, progress, and important decisions outside the model so a fresh session can reconstruct context deterministically
- prefer tools and checks that let the system verify itself instead of asking the model to self-certify completion
- treat noisy, stale, or weakly relevant context as a design bug
If you remember only one rule, remember this:
- always construct the cleanest, most task-relevant, shortest complete context the system can maintain in a closed loop
Best-practice patterns
Instant shell-first usage (Recommended for one-offs)
Use the direct constructor path when you want to turn SimpleLLMFunc into a shell ability with almost no setup besides pasting your literal model settings.
python - <<'PY'
import asyncio
from SimpleLLMFunc import APIKeyPool, OpenAICompatible, llm_function
llm = OpenAICompatible(
api_key_pool=APIKeyPool(
api_keys=["sk-your-key"],
provider_id="openrouter-z-ai-glm-5",
),
model_name="z-ai/glm-5",
base_url="https://openrouter.ai/api/v1",
)
@llm_function(llm_interface=llm)
async def answer(question: str) -> str:
"""Answer the question in a compact, practical way."""
pass
print(asyncio.run(answer("Give me three uses of SimpleLLMFunc.")))
PY
For a shell agent with file tools and REPL, see reference/instant-use.md and examples/instant_chat_agent.py.
Typed llm_function (Recommended)
import asyncio
from pydantic import BaseModel, Field
from SimpleLLMFunc import OpenAICompatible, llm_function
class SentimentReport(BaseModel):
sentiment: str = Field(description="positive, negative, or neutral")
confidence: float = Field(description="0.0 to 1.0 confidence score")
summary: str = Field(description="one-sentence explanation")
models = OpenAICompatible.load_from_json_file("provider.json")
llm = models["openrouter"]["z-ai/glm-5"]
@llm_function(llm_interface=llm)
async def classify_sentiment(text: str) -> SentimentReport:
"""
Classify the sentiment of the input text.
Args:
text: The user text to analyze.
Returns:
A structured sentiment report.
"""
pass
async def main() -> None:
result = await classify_sentiment("The setup was rough, but the product is excellent.")
print(result.model_dump())
asyncio.run(main())
Chat agent with a tool
import asyncio
from SimpleLLMFunc import OpenAICompatible, llm_chat, tool
@tool
async def multiply(a: float, b: float) -> float:
"""
Multiply two numbers.
Args:
a: First factor.
b: Second factor.
Returns:
Product of the two inputs.
"""
return a * b
models = OpenAICompatible.load_from_json_file("provider.json")
llm = models["openrouter"]["z-ai/glm-5"]
@llm_chat(llm_interface=llm, toolkit=[multiply], stream=True)
async def tutor(message: str, history: list[dict[str, str]] | None = None):
"""
You are a concise math tutor.
Use the multiply tool when arithmetic is requested.
"""
pass
async def main() -> None:
history: list[dict[str, str]] = []
async for chunk, history in tutor("What is 12.5 times 8?", history):
if chunk:
print(chunk, end="")
asyncio.run(main())
Persistent runtime with PyRepl
import asyncio
from SimpleLLMFunc.builtin import PyRepl, SelfReference
MEMORY_KEY = "agent_main"
async def main() -> None:
repl = PyRepl()
selfref = repl.get_runtime_backend("selfref")
assert isinstance(selfref, SelfReference)
selfref.bind_history(
MEMORY_KEY,
[{"role": "system", "content": "Answer in bullet points."}],
)
result = await repl.execute(
"runtime.selfref.context.remember('remember this')\n"
"snapshot = runtime.selfref.context.inspect()\n"
"print(len(snapshot['experiences']))"
)
print(result["stdout"])
asyncio.run(main())
Hard rules and gotchas
- Default to
async deffor all decorators.@toolenforces async directly. - The function body does not implement behavior; the docstring does.
- The docstring is usually only the base material for prompt construction, not the entire final system prompt.
- For
llm_chat, name the history parameterhistoryorchat_history. - The current direct-construction path for
OpenAICompatiblerequiresAPIKeyPool; do not assume a simplifiedapi_key=constructor exists. - Instant snippets can call the decorated function directly at top level with
asyncio.run(...); no__main__guard is required. max_tool_calls=Nonemeans no framework-imposed tool-call cap. Set an explicit integer if you need a guardrail.- Complex structured outputs are parsed from XML-oriented contracts internally. Do not manually force JSON unless you intentionally want plain-text behavior.
llm_chat(enable_event=True)yieldsReactOutputevents and responses, not(chunk, history)tuples.too_long_to_file=Truekeeps roughly the first 20000 tokens in chat and writes the full tool result to a temp file.PyRepl.reset()clears REPL variables but keeps runtime backends and self-reference memory.runtime.selfref.context.compact(...)is queued first. When called from a tool run, the compacted context is applied before the next same-turn LLM step when possible, and finalize still commits any leftover queued compaction before the turn ends.OpenAIResponsesCompatibleis a first-class adapter. It maps the selected system prompt to Responsesinstructions, supportsreasoning={...}, and keeps Responses-specific request/stream behavior out of your decorator code.runtime.selfref.fork.spawn(...)children inherit the pre-fork context snapshot, not the parent's in-flight fork tool-call scene.runtime.selfref.fork.gather_all(...)returnsdict[fork_id -> ForkResult]. Checkstatusfirst, then readresponseorresult; compact results omit child history unless you requestinclude_history=True.FileToolsetis workspace-scoped and read-before-write guarded.
Load more context only when needed
- Philosophy and core concepts:
reference/philosophy-and-concepts.md - Harness engineering guidance:
reference/harness-engineering.md - System prompt construction and prompt-writing rules:
reference/system-prompt-construction.md - Instant shell-first setup and constructor usage:
reference/instant-use.md - Provider and environment setup:
reference/configuration.md - Decorators, tools, file tools, and event mode:
reference/decorators-and-tools.md - PyRepl, runtime primitives, and selfref:
reference/pyrepl-runtime.md - Non-obvious behavior:
reference/gotchas.md - Mirrored repo docs:
reference/docs-source/quickstart.md,reference/docs-source/guide.md,reference/docs-source/detailed_guide/ - Instant heredoc examples:
examples/instant_llm_function.py,examples/instant_chat_agent.py - Real repo examples:
examples/agent_as_tool_example.py,examples/llm_function_pydantic_example.py,examples/runtime_primitives_basic_example.py,examples/tui_general_agent_example.py