文件历史

328 次代码提交

作者 SHA1 备注 提交日期
Simon Willison 527f84fd5d DAG phase 6: llm logs surfaces DAG head pointers via calls join
LOGS_SQL and LOGS_SQL_SEARCH left join the calls table on responses.id
(they share ids by construction). --json output now includes
head_input_message_id and head_output_message_id.

Rows that predate the DAG schema — or historical fixtures that only
wrote to responses — get NULL for those columns, so existing tests
are unaffected. Two new tests cover both cases: real prompt round-trip
(fields populated) and legacy-shape rows (fields NULL).

This is the minimum hook from the Deferred list in plans/dag-schema.md.
Union-read (prefer calls over responses, reconstruct prompt/response
from the DAG) is still deferred — a follow-up once there are calls
rows without matching responses rows.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-04-14 08:34:35 -07:00
Simon Willison 94cde01b30 DAG phase 4: fork
- MessageStore.fork(source_message_id, name=, model=) creates a new
  conversation rooted at an existing message. The shared prefix is
  reused in place; only one conversations row is written.
- Model defaults to whatever most recently wrote a call touching the
  source message; can be overridden. Raises ValueError if the message
  doesn't exist or no model can be inferred.
- New `llm fork <message_id>` CLI prints the new conversation id.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-04-14 08:27:38 -07:00
Simon Willison 1985f98401 DAG phase 3: continuation dedup + Conversation.head_message_id
- Conversation (sync + async) gains head_message_id; from_row populates
  it from the DB so CLI -c continuation resumes the existing chain
  instead of starting a parallel one.
- MessageStore gains save_with_dedup(): one call that locates the
  longest existing prefix and appends only the unmatched tail. Writes
  zero rows when the full chain already exists — the stateless-API
  continuation case from plans/dag-schema.md.
- Tests cover find_longest_existing_prefix (miss / partial / full),
  save_with_dedup idempotency, and end-to-end: a second prompt on a
  reloaded Conversation extends the chain via head_message_id.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-04-14 08:26:13 -07:00
Simon Willison ee08b572c8 DAG phase 2: schema replacement + MessageStore
Rewrites m023 in place to the DAG-shaped message store from
plans/dag-schema.md:

- messages: id, parent_id, content_hash, role, provider_metadata_json,
  created_at. Chain roots point at a self-referencing sentinel row
  ("root") so the unique (parent_id, content_hash) index works at
  every chain position — NULL-parent uniqueness footgun avoided.
- message_parts: structurally unchanged.
- calls: one row per LLM call, anchoring head_input/head_output
  message ids and recording model + timing + usage.
- conversations.head_message_id: advances each turn; history is
  reconstructed by walking parent_id from the head.

New llm/storage.py provides MessageStore.save_chain (with dedup),
load_chain, and find_longest_existing_prefix (for the stateless-API
case wired in phase 3).

Response.log_to_db now writes the DAG + a calls row alongside the
existing responses-table writes (kept for llm logs compatibility
until phase 5). Response._load_messages_from_db walks the chain
using calls pointers.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-04-14 08:23:31 -07:00
Simon Willison 3bf5606b5a DAG phase 1: canonical message serialization + content hash
Adds llm/_canonical.py with canonical_message_json() and
message_content_hash(). The serialization is the hash contract for the
incoming DAG-shaped message store; snapshot tests in
tests/test_canonical.py pin the wire format.

The include_provider_metadata parameter is threaded through now so a
future semantic_hash column can be added without refactoring
(see plans/dag-provider-metadata-hashing.md).

Also commits the design docs:
- plans/dag-schema.md — full DAG storage design
- plans/dag-provider-metadata-hashing.md — Option 1 now, Option 2 later

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-04-14 08:12:45 -07:00
Simon Willison d6e810226a Remove role attribute from Part classes
Role now lives exclusively on the enclosing Message. Part subclasses no
longer accept or expose role; to_dict / from_dict drop the role key.
normalize_parts() lost its role= parameter (and helpers stopped passing
one), since emitted parts inherit their role from the Message they're
constructed into.

_parts_to_messages now wraps an output parts list in a single
assistant Message (which matches how providers like Anthropic package
server-side tool results — inside the assistant turn's content
blocks). Multi-message responses are represented via the
message_index field on StreamEvent (added earlier), not by reading
role off individual parts.

Tests updated to drop part.role assertions and to_dict role keys.
2026-04-12 21:48:25 -07:00
Simon Willison d34449a503 DB: new messages + message_parts schema
Migration m023 creates messages and message_parts tables. The old
parts table (from m022) is left in place for databases that already
ran that migration but is no longer read or written.

log_to_db now walks prompt.messages and response.messages, inserting
one messages row per Message and one message_parts row per Part.
Response.from_row loads via _load_messages_from_db into
_loaded_messages, which Response.messages returns directly — no more
group-parts-back-into-messages dance on load.

Tests updated to assert against the new schema. Per the branch
decision to ignore prior logs, no backfill migration is provided.
2026-04-12 21:40:00 -07:00
Simon Willison 9c4ac33b73 Remove parts=[] parameter and response.parts property
Hard removal of the flat parts-as-input API. Callers now use messages=
(with user/assistant/system/tool_message helpers) for explicit history
and response.messages for the structured response.

Changes:
- Prompt drops _parts and the parts= kwarg; Prompt.parts property gone.
- Response.parts and AsyncResponse.parts properties gone; messages is
  now the canonical accessor.
- Model.prompt / Conversation.prompt / async variants drop parts= kwarg.
- log_to_db walks self.prompt.messages for input rows and
  self._build_parts() for output rows, via a shared _part_to_row helper.

Tests migrated: TestPartsParameter deleted, TestBuildMessagesWithParts
rewritten to messages=. All remaining response.parts / r.parts /
loaded.parts usages flattened over response.messages.
2026-04-12 21:35:20 -07:00
Simon Willison aae4eedf16 OpenAI: reconstruct conversation input history via prev.prompt.messages
Iterate prev_response.prompt.messages (the computed Message list) to
rebuild prior turn inputs, and funnel each through
_append_message_from_message. Output side still uses the flat text /
tool_calls accumulators (text_or_raise, tool_calls_or_raise) to avoid
calling _build_parts on historical responses whose StreamEvent shape
might have used the same part_index for mixed content types.
2026-04-12 21:20:57 -07:00
Simon Willison 6f961653ef Add Response.messages (additive)
response.messages groups the flat parts list into a list of Message
objects by consecutive role. Typical case: one assistant message
wrapping all parts. Server-executed tool results (role='tool')
interleaved among assistant parts produce additional Messages at role
boundaries.

response.parts still works; hard-removal is deferred to a later commit
so mechanical test migration can happen in one focused change.
2026-04-12 21:16:10 -07:00
Simon Willison a6a25b26e5 Make Prompt.messages authoritative, OpenAI adapter reads only messages
prompt.messages is now a computed property that synthesizes Messages
from legacy inputs (system=, parts=, prompt=, attachments=,
tool_results=) when messages= was not explicitly passed. Explicit
messages= passes through verbatim.

OpenAI build_messages() for the current prompt now has a single code
path that iterates prompt.messages — the old if/elif over _parts vs
legacy fields is gone. Conversation history reconstruction still uses
legacy fields (will flip in a later commit).
2026-04-12 21:13:42 -07:00
Simon Willison d833616276 Add messages=[...] parameter to model.prompt()
Accept messages= alongside the existing parts= parameter.
Conversation/AsyncConversation/Model/AsyncModel prompt() forward it to
Prompt, which stores it as prompt.messages.

OpenAI adapter gains _append_message_from_message which translates one
llm.Message into the correct OpenAI message dict(s), including the
parallel-tool-calls case (one assistant message with multiple
ToolCallParts becomes one OpenAI message with a tool_calls array).

Legacy paths untouched: parts=, prompt=, system=, attachments=, and
tool_results= still work when messages= is not set.
2026-04-12 21:10:28 -07:00
Simon Willison f4e955705f Add Message class and role helpers (user/assistant/system/tool_message)
First step toward replacing flat parts=[] with structured messages=[].
Introduces the Message dataclass (role + parts + provider_metadata) and
convenience helpers that normalize strings, Attachments, Parts, and
nested lists into Message objects.

Additive only: existing Part.role, response.parts, parts= API still
work. Later commits flip Prompt/Response to consume messages and
remove the legacy surface.
2026-04-12 21:07:12 -07:00
Simon Willison 5217a0bace Parts: serialization round-trip, stream safety, provider_metadata
Round out the parts API so transcripts survive serialization and so
providers can stash opaque multi-turn state on parts and stream events.

Serialization:
- AttachmentPart.from_dict now supported; inline content bytes round-trip
  as base64.
- ToolResultPart.attachments round-trip through to_dict/from_dict.

Stream assembler:
- _build_parts raises ValueError when an incompatible StreamEvent type
  appears at the same part_index, instead of silently overwriting the
  earlier part. tool_call_name and tool_call_args stay compatible.

OpenAI parts=[] support:
- build_messages emits assistant tool_calls and role:"tool" messages for
  ToolCallPart and ToolResultPart passed via parts=.

provider_metadata:
- New optional dict on TextPart, ReasoningPart, ToolCallPart,
  ToolResultPart, and StreamEvent for opaque provider data that must be
  echoed back on the next request (Anthropic signature/encrypted_content,
  Gemini thoughtSignature, OpenAI Responses encrypted_content).
- StreamEvent values merge onto the finalized Part per top-level namespace
  key, last non-None wins.
- Persisted via existing content_json column and reloaded by
  _load_parts_from_db; no schema change.
- Plugin author guide in docs/plugins/advanced-model-plugins.md.

Types:
- Widen execute() return types to Iterator[str | StreamEvent] /
  AsyncGenerator[str | StreamEvent, None] on abstract Model/AsyncModel
  bases and OpenAI Chat implementations.
- Initialize _reasoning_token_count on _BaseResponse so mypy stops
  flagging the OpenAI plugin.
2026-04-12 20:59:21 -07:00
Simon Willison f773ab2fa4 Formatting and type annotation fixes 2026-04-12 17:31:21 -07:00
Simon Willison f308149dee Fix some ruff linter warnings 2026-04-12 17:21:04 -07:00
Simon Willison 7c0a192341 Reformatted with Black
Test / test (macos-latest, 3.10) (push) Has been cancelled
Test / test (macos-latest, 3.11) (push) Has been cancelled
Test / test (macos-latest, 3.12) (push) Has been cancelled
Test / test (macos-latest, 3.13) (push) Has been cancelled
Test / test (macos-latest, 3.14) (push) Has been cancelled
Test / test (ubuntu-latest, 3.10) (push) Has been cancelled
Test / test (ubuntu-latest, 3.11) (push) Has been cancelled
Test / test (ubuntu-latest, 3.12) (push) Has been cancelled
Test / test (ubuntu-latest, 3.13) (push) Has been cancelled
Test / test (ubuntu-latest, 3.14) (push) Has been cancelled
Test / test (windows-latest, 3.10) (push) Has been cancelled
Test / test (windows-latest, 3.11) (push) Has been cancelled
Test / test (windows-latest, 3.12) (push) Has been cancelled
Test / test (windows-latest, 3.13) (push) Has been cancelled
Test / test (windows-latest, 3.14) (push) Has been cancelled
2026-04-12 17:18:37 -07:00
Simon Willison 1280082b6d Rename prompt.input_parts to prompt.parts 2026-04-12 17:17:42 -07:00
Simon Willison 1610030020 Default OpenAI plugin now handles .prompt(parts=[...]) 2026-04-07 12:18:51 -07:00
Simon Willison 018217d624 Add newline at reasoning-to-text transition, extract display helper
display_stream_events() helper handles writing text to stdout and
reasoning to stderr with proper newlines at each reasoning-to-text
transition. Used by sync prompt and chat streaming loops. Async prompt
loop has the same logic inline.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-04-05 07:59:54 -07:00
Simon Willison 05f5ed3943 CLI displays reasoning on stderr, adds -R/-S/-L shortcuts
Streaming loops in prompt and chat commands now use stream_events()
instead of __iter__. Reasoning events are displayed on stderr in
dim text. Text events go to stdout as before.

New flags:
  -R / --no-reasoning  Suppress reasoning output on stderr
  -S / --no-stream     (shortcut for existing --no-stream)
  -L                   (shortcut for existing -n/--no-log)

ChainResponse.stream_events() and AsyncChainResponse.astream_events()
added so tool-calling flows also surface reasoning events.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-04-05 07:53:25 -07:00
Simon Willison 0205bec5d1 Database migration and persistence for parts (Phase 5)
New m022_parts_table migration creates a parts table with direction
(input/output), role, part_type, content, content_json, tool_call_id,
and server_executed columns.

log_to_db() writes both input parts (from prompt.input_parts) and
output parts (from response.parts) to the table. from_row() loads
output parts and makes them available via the parts property.

Tested live: parts table created, input/output parts written and
loaded correctly with gpt-5.4-mini via CLI.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-04-05 07:42:26 -07:00
Simon Willison 5194bb766f Add parts= parameter and input_parts property (Phase 4)
model.prompt() and conversation.prompt() now accept a parts= parameter
for passing explicit Part objects. Prompt.input_parts synthesizes a
unified list of input Parts from prompt=, system=, attachments=, and
parts= parameters.

prompt= remains sugar for a TextPart(role="user"). system= becomes
a TextPart(role="system"). attachments= become AttachmentParts.
All parameters combine (parts first, then system, then prompt, then
attachments).

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-04-05 07:38:45 -07:00
Simon Willison 0367001d18 OpenAI plugin emits StreamEvent objects (Phase 3)
Chat.execute() and AsyncChat.execute() now yield StreamEvent instead
of bare strings. Text chunks, tool call names/args are all emitted
as typed events. Reasoning token counts from usage data are stored
on the response and _build_parts() prepends a redacted ReasoningPart.

StreamEvent gains server_executed and tool_name fields for use by
plugins with server-side tool execution.

Tested live against gpt-5.4-mini: text streaming, tool calls, and
reasoning tokens (with reasoning_effort='high') all work correctly.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-04-05 07:24:40 -07:00
Simon Willison a8e7b63604 Teach Response to handle StreamEvent from plugins (Phase 2)
Response.__iter__ now handles str | StreamEvent from execute().
Plain str yields are backward compatible. StreamEvent yields are
processed by the assembler: text events yield as str to consumers,
reasoning/tool_call/tool_result events are filtered from __iter__
but available via stream_events(). Parts are assembled from events
after completion. Same changes for AsyncResponse.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-04-04 23:23:22 -07:00
Simon Willison c774a22ab9 Export Part types and StreamEvent from llm package
All new types are now accessible as llm.TextPart, llm.StreamEvent, etc.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-04-04 23:20:10 -07:00
Simon Willison 350469f266 Add stream_events() and parts property to Response/AsyncResponse
Response.stream_events() yields StreamEvents wrapping text chunks.
AsyncResponse.astream_events() is the async equivalent.
response.parts returns a list of Part objects after completion.
Currently only handles plain str chunks (Phase 1 baseline).

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-04-04 23:18:57 -07:00
Simon Willison af1311de23 Part types and StreamEvent dataclasses with serialization
Phase 1 of the parts project: define the Part dataclass hierarchy
(TextPart, ReasoningPart, ToolCallPart, ToolResultPart, AttachmentPart)
and StreamEvent in a new llm/parts.py module. All Part types have
to_dict()/from_dict() for JSON roundtripping. 14 tests.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-04-04 23:17:34 -07:00
Simon Willison cad03fb4f4 Register async models for extra-openai-models.yaml, closes #1395
Note that Completion models do not have an async class so will not be registered as async.
2026-04-04 07:07:14 -07:00
Simon Willison c8889e0a76 Ran black 2026-03-17 11:25:12 -07:00
Simon Willison 683ca204b2 Ensure -x/--xl work with -t 2026-03-17 11:22:34 -07:00
Simon Willison 5d237ce6ce Show options in Markdown logs output, closes #1322 2025-12-17 22:19:33 -08:00
Simon Willison e0b44bc5ab Fix some test warnings, refs #1312 2025-12-11 14:37:47 -08:00
Eric Bloch f7934c5c26 Fix some descriptor leaks (#1313)
Refs #1312
2025-12-11 14:20:58 -08:00
Simon Willison a0ac68c452 Custom HTTP user-agent for llm -f URL, closes #1309 2025-11-25 22:07:11 -08:00
Simon Willison a04a6afa74 AsyncModel in llm.__all__, closes #1308 2025-11-25 13:44:42 -08:00
Simon Willison c41c122239 Use tools in templates with llm chat, closes #1239 2025-08-11 22:11:17 -07:00
Simon Willison 2f206d0e26 Fix for duplicated prompts in llm chat with templates, closes #1240
Also includes a bug fix for system fragments, see https://github.com/simonw/llm/issues/1240#issuecomment-3177684518
2025-08-11 21:52:54 -07:00
Simon Willison c6e158071a Ran black 2025-08-11 21:47:11 -07:00
Simon Willison e6ac18fbcb Fix for confusing error, closes #1238 2025-08-11 16:52:16 -07:00
Simon Willison 9f1417f6e8 Fix for enum options and --save, refs #1237 2025-08-11 16:16:27 -07:00
Simon Willison e4c1a46d90 Fix test failure caused by version bump, refs #1218 2025-08-11 14:27:16 -07:00
James Sanford 2a54939951 Fix streaming tool calls with tests for many variants. (#1218)
* Recorded instance of streaming tool response variant "a".

This is the typical response, where "arguments":"" arrives
first in the stream, followed by "arguments":"{}"

The response data is a real capture from the OpenRouter API,
however some request and header data may be from other test fixtures.

* Recorded instance of streaming tool response variant "b".

This is a streaming response where the first arguments
you get is a fully formed "arguments":"{}"

The response data is a real capture from the OpenRouter API,
however some request and header data may be from other test fixtures.

* Test cases for streaming tool responses.

Note that the replays are marked as "read-only", as they are variants
seen in the wild where the streaming tool call argument fragments
arrive in a specific order.

* Fix streaming tool response variant "b", where "arguments":"{}" is what arrives first.

The previous code erroneously caused the first "arguments" to be duplicated,
by using "+=" even when being initially set.

This went unnoticed as many models stream "arguments":"" first.

When a more fully formed "arguments" fragment arrived first, it was causing
"Error: Extra data: line 1 column 3 (char 2)"

* Recorded instance of streaming tool response variant "c".

This was failing with "Error: unsupported operand type(s) for +=: 'NoneType' and 'str'"

The response data is a real capture from the OpenRouter API,
however some request and header data may be from other test fixtures.

* Test case for streaming tool response variant "c".

* Fix streaming tool response variant "c".

However, I'm not sure why arguments was initially not present or seen as None.
2025-08-11 14:17:52 -07:00
Simon Willison 5204a11f33 Allow -o option when calling tools, closes #1233 2025-08-11 13:44:26 -07:00
Simon Willison 08094082f2 Toolbox.add_tool(), prepare() and prepare_async() methods
Closes #1111
2025-08-11 13:19:31 -07:00
Simon Willison 0863ed460e llm logs -l/--latest -q option, closes #1177 2025-06-17 23:23:45 -07:00
Simon Willison 544ce17c1d Tests to confirm responses FTS triggers
Refs https://github.com/simonw/llm/issues/1177#issuecomment-2982832935
2025-06-17 23:18:22 -07:00
Simon Willison 3a96d52895 Better handling of before_call cancellation, closes #1148 2025-06-01 18:36:55 -07:00
Simon Willison d96ae4ed8d Fix --async logging to database, closes #1150 2025-06-01 17:38:26 -07:00
Simon Willison 30e0c4abe8 ToolResult.exception for tool errors, now logged to DB
Closes #1104
2025-06-01 17:01:40 -07:00