文件历史

135 次代码提交

作者 SHA1 备注 提交日期
Ali Khokhar 2a676cc6d9 Make ProviderModelInfo the sole model-catalog contract (#1222)
## Problem

Provider model discovery consumes metadata, but providers and the cache
still expose a parallel IDs-only contract. The duplicate contract adds
adapters and lets tests bypass capability metadata.

## Changes

| Before | After |
| --- | --- |
| `BaseProvider` exposed `list_model_ids()` plus a metadata adapter. |
`BaseProvider` exposes only abstract `list_model_infos()` returning
application-owned metadata. |
| Ordinary providers parsed IDs and converted them later. | Ordinary
providers parse OpenAI-compatible catalogs directly into
`ProviderModelInfo` values. |
| OpenRouter, Cloudflare, and GitHub Models maintained redundant
IDs-only wrappers. | Provider-specific filters and capability metadata
have one return path. |
| Vertex returned paginated IDs for the base adapter to wrap. | Vertex
returns metadata after completing the same paginated discovery flow. |
| The runtime cache exposed test-only raw-ID write and prefixed-ID read
helpers. | The runtime cache accepts and returns metadata while
retaining its production admin-status ID projection. |
| Provider tests asserted the parallel IDs-only API. | Provider tests
enforce the metadata-only contract and preserve provider-specific
discovery behavior. |
| The package version was `4.11.6`. | The package version is `4.11.7`
with an updated lockfile. |

<!-- greptile_comment -->

<details open><summary><h3>Greptile Summary</h3></summary>

This PR makes provider metadata the only model-catalog contract. The
main changes are:

- Makes `list_model_infos()` the abstract provider discovery API.
- Migrates provider parsers and implementations to `ProviderModelInfo`.
- Removes IDs-only cache and parser helpers.
- Updates provider tests, architecture documentation, and package
metadata.
</details>

<h3>Confidence Score: 4/5</h3>

The catalog migration is consistent, but the release version must
reflect the incompatible API removal.

Repository provider and cache call paths use the new metadata shape
consistently. Existing external consumers of the removed contracts can
fail after a patch upgrade. The repository rules classify incompatible
API removals as a major release.

pyproject.toml and the matching package entry in uv.lock

<details><summary><h3><a href="https://www.greptile.com/trex"><img
alt="T-Rex"
src="https://greptile-static-assets.s3.amazonaws.com/trex/trex_green.svg"
height="20" align="absmiddle"></a> T-Rex Logs</h3></summary>

**What T-Rex did**
- The external-consumer compatibility probe was run against both
revisions, confirming the base still supports the legacy provider
contract, while head fails with a TypeError due to the missing
list\_model\_infos method, and runtime metadata shows head at version
4.11.7, indicating the removal is a patch transition rather than a major
change.
- An automated test suite completed successfully with 83 tests passing
in 3.31 seconds.

<a
href="https://app.greptile.com/trex/runs/15207030/artifacts"><picture><source
media="(prefers-color-scheme: dark)"
srcset="https://greptile-static-assets.s3.amazonaws.com/badges/ViewAllArtifactsDark.svg?v=4"><source
media="(prefers-color-scheme: light)"
srcset="https://greptile-static-assets.s3.amazonaws.com/badges/ViewAllArtifacts.svg?v=4"><img
alt="View all artifacts"
src="https://greptile-static-assets.s3.amazonaws.com/badges/ViewAllArtifacts.svg?v=4"></picture></a>

<sub><a href="https://www.greptile.com/trex"><img alt="T-Rex"
src="https://greptile-static-assets.s3.amazonaws.com/trex/trex_green.svg"
height="14" align="absmiddle"></a> Ran code and verified through
T-Rex</sub>
</details>

<details open><summary><h3>Important Files Changed</h3></summary>

| Filename | Overview |
|----------|----------|
| src/free_claude_code/providers/base.py | Replaces the IDs-only
provider API with an abstract metadata-only contract. |
| src/free_claude_code/providers/model_listing.py | Consolidates
OpenAI-compatible parsing into ProviderModelInfo results and removes
IDs-only helpers. |
| src/free_claude_code/providers/runtime/model_cache.py | Removes raw-ID
helpers while retaining metadata storage and the admin ID projection. |
| src/free_claude_code/providers/openai_chat/provider.py | Provides the
metadata discovery implementation inherited by ordinary
OpenAI-compatible providers. |
| src/free_claude_code/providers/vertex/client.py | Preserves paginated
discovery while returning metadata values. |
| pyproject.toml | Uses a patch bump for a release that removes callable
and importable contracts. |

</details>

<details open><summary><h3>Flowchart</h3></summary>

<a href="#gh-light-mode-only">

```mermaid
%%{init: {'theme': 'neutral'}}%%
flowchart LR
A[Provider catalog endpoint] --> B[list_model_infos]
B --> C[ProviderModelInfo set]
C --> D[Provider model discovery]
D --> E[ProviderModelCache]
E --> F[Metadata-aware catalog]
E --> G[Admin ID projection]
```

</a>
<a href="#gh-dark-mode-only">

```mermaid
%%{init: {'theme': 'base', 'themeVariables': {"darkMode": true, "background": "#0d1117", "primaryColor": "#21262d", "primaryTextColor": "#e6edf3", "primaryBorderColor": "#8b949e", "lineColor": "#8b949e", "textColor": "#e6edf3", "edgeLabelBackground": "#161b22", "actorBkg": "#21262d", "actorBorder": "#8b949e", "actorTextColor": "#e6edf3", "actorLineColor": "#8b949e", "signalColor": "#8b949e", "signalTextColor": "#e6edf3", "noteBkgColor": "#373320", "noteBorderColor": "#d4a72c", "noteTextColor": "#f0e6c0", "labelBoxBkgColor": "#21262d", "labelBoxBorderColor": "#8b949e", "labelTextColor": "#e6edf3", "loopTextColor": "#e6edf3", "activationBkgColor": "#30363d", "activationBorderColor": "#8b949e"}}}%%
flowchart LR
A[Provider catalog endpoint] --> B[list_model_infos]
B --> C[ProviderModelInfo set]
C --> D[Provider model discovery]
D --> E[ProviderModelCache]
E --> F[Metadata-aware catalog]
E --> G[Admin ID projection]
```

</a>
</details>

<a
href="https://app.greptile.com/api/ide/codex?prompt=IMPORTANT%3A%20Work%20in%20the%20repository%20%22alishahryar1%2Ffree-claude-code%22%20on%20the%20existing%20branch%20%22ali%2Fprovider-model-info-contract%22.%20Checkout%20that%20branch%20%E2%80%94%20do%20NOT%20create%20a%20new%20branch%20or%20open%20a%20new%20PR.%20Push%20your%20changes%20to%20%22ali%2Fprovider-model-info-contract%22.%0A%0AFix%20the%20following%201%20code%20review%20issue.%20Work%20through%20them%20one%20at%20a%20time%2C%20proposing%20concise%20fixes.%0A%0A---%0A%0A%23%23%23%20Issue%201%20of%201%0Apyproject.toml%3A7%0A**Breaking%20Contract%20Ships%20as%20Patch**%0A%0AThis%20release%20removes%20%60BaseProvider.list_model_ids%28%29%60%20and%20cache%2Fparser%20methods%20that%20existing%20integrations%20can%20import%20or%20call.%20Such%20consumers%20will%20fail%20with%20%60TypeError%60%2C%20%60AttributeError%60%2C%20or%20%60ImportError%60%20after%20a%20patch%20upgrade%2C%20so%20this%20incompatible%20API%20change%20requires%20a%20major%20version%20bump%20under%20the%20repository's%20versioning%20rules.%0A%0A%60%60%60suggestion%0Aversion%20%3D%20%225.0.0%22%0A%60%60%60%0A%0A&repo=alishahryar1%2Ffree-claude-code&pr=1222&platform=github"><picture><source
media="(prefers-color-scheme: dark)"
srcset="https://greptile-static-assets.s3.amazonaws.com/badges/FixAllInCodexDark.svg?v=6"><source
media="(prefers-color-scheme: light)"
srcset="https://greptile-static-assets.s3.amazonaws.com/badges/FixAllInCodex.svg?v=6"><img
alt="Fix All in Codex"
src="https://greptile-static-assets.s3.amazonaws.com/badges/FixAllInCodex.svg?v=6"></picture></a>

<sub>Reviews (1): Last reviewed commit: ["Make ProviderModelInfo the
sole catalog
..."](https://github.com/alishahryar1/free-claude-code/commit/5c543eadc114201a4085d38885a27ac49153ca29)
| [Re-trigger
Greptile](https://app.greptile.com/api/retrigger?id=45869248)</sub>

> Greptile also left **1 inline comment** on this PR.

**Context used:**

- Context used - CLAUDE.md
([source](https://app.greptile.com/alishahryar1/github/Alishahryar1/free-claude-code/-/custom-context?memory=d2fd24d8-0dec-4faf-8ee4-e085e215a2f8))

<!-- /greptile_comment -->
2026-07-21 00:15:35 -07:00
Ali Khokhar 1a476562fc Make provider model discovery the sole catalog owner (#1221)
## Problem

Startup model-list I/O was split between a non-enforcing
configured-model validator and the real discovery path. Both queried
providers and populated the same cache even though synchronous cache
warm-up was the validator's only required effect.

## Changes

| Before | After |
| --- | --- |
| Validation and discovery independently resolved providers, queried
model lists, and cached results. | `ProviderModelDiscovery` solely owns
model-list queries, failure reporting, and cache population. |
| Startup ran configured-model validation and then launched discovery. |
Startup synchronously warms referenced providers through discovery, then
launches the existing missing-provider background pass. |
| Absence from a provider catalog produced a non-enforcing missing-model
warning. | Provider catalogs remain discovery metadata and provider
execution remains authoritative. |
| Configured model references retained environment-source metadata for
validator diagnostics. | Configured model references retain only routing
data. |
| Successful startup providers could be represented by two separate
subsystems. | Focused tests enforce concurrent warm-up, single
successful queries, failed-query eligibility, and warm-before-background
ordering. |
| The package version was `4.11.5`. | The package version is `4.11.6`
with an updated lockfile. |

<!-- greptile_comment -->

<details open><summary><h3>Greptile Summary</h3></summary>

This PR makes provider discovery the sole owner of model catalogs. The
main changes are:

- Warms routed provider catalogs before background discovery starts.
- Reuses successful warm results while retrying failed providers.
- Removes configured-model catalog validation and source metadata.
- Moves query-failure reporting into the discovery module.
- Updates focused tests, architecture docs, and package version.
</details>

<h3>Confidence Score: 5/5</h3>

This looks safe to merge.

No blocking issues found in the changed code. Provider failures remain
isolated and eligible for background retry. Successful warm results are
not queried again by the missing-provider pass. Lease release and
startup cleanup remain protected by existing control flow.

None.

<details><summary><h3><a href="https://www.greptile.com/trex"><img
alt="T-Rex"
src="https://greptile-static-assets.s3.amazonaws.com/trex/trex_green.svg"
height="20" align="absmiddle"></a> T-Rex Logs</h3></summary>

**What T-Rex did**
- \`test\_runtime\_warm\_queries\_referenced\_providers\_concurrently\`,
\`test\_startup\_discovery\_queries\_each\_successful\_provider\_once\`,
\`test\_failed\_startup\_warm\_remains\_eligible\_for\_background\_refresh\`,
\`test\_runtime\_warm\_caches\_all\_referenced\_provider\_models\`, and
\`test\_runtime\_startup\_warms\_catalog\_before\_background\_refresh\`
all passed.
- The complete verbose HEAD run, including command, working directory,
exit code, test nodes, and summary, is preserved in
\`trex-artifacts/provider-discovery-startup-validation.log\` and the
paired after artifact.

<a
href="https://app.greptile.com/trex/runs/15203372/artifacts"><picture><source
media="(prefers-color-scheme: dark)"
srcset="https://greptile-static-assets.s3.amazonaws.com/badges/ViewAllArtifactsDark.svg?v=4"><source
media="(prefers-color-scheme: light)"
srcset="https://greptile-static-assets.s3.amazonaws.com/badges/ViewAllArtifacts.svg?v=4"><img
alt="View all artifacts"
src="https://greptile-static-assets.s3.amazonaws.com/badges/ViewAllArtifacts.svg?v=4"></picture></a>

<sub><a href="https://www.greptile.com/trex"><img alt="T-Rex"
src="https://greptile-static-assets.s3.amazonaws.com/trex/trex_green.svg"
height="14" align="absmiddle"></a> Ran code and verified through
T-Rex</sub>
</details>

<details open><summary><h3>Important Files Changed</h3></summary>

| Filename | Overview |
|----------|----------|
| src/free_claude_code/providers/runtime/discovery.py | Centralizes
catalog queries, failure reporting, referenced-provider warming, and
cache population. |
| src/free_claude_code/runtime/provider_manager.py | Replaces validation
with discovery-based warming under a generation lease. |
| src/free_claude_code/runtime/application.py | Warms referenced
catalogs before launching the missing-provider background pass. |
| src/free_claude_code/config/model_refs.py | Removes validation-only
source metadata while preserving deterministic deduplication. |
| tests/providers/test_model_discovery.py | Covers concurrent warming,
partial failures, retry eligibility, and single successful queries. |

</details>

<details open><summary><h3>Sequence Diagram</h3></summary>

<a href="#gh-light-mode-only">

```mermaid
%%{init: {'theme': 'neutral'}}%%
sequenceDiagram
participant App as ApplicationRuntime
participant Manager as ProviderRuntimeManager
participant Discovery as ProviderModelDiscovery
participant Provider
participant Cache as ProviderModelCache

App->>Manager: warm_referenced_model_cache()
Manager->>Discovery: warm referenced providers
par Provider queries
    Discovery->>Provider: list_model_infos()
end
Provider-->>Discovery: metadata or failure
Discovery->>Cache: cache successful results
Discovery-->>Manager: refresh result
Manager-->>App: warm complete
App->>Manager: start_model_list_refresh()
Manager->>Discovery: refresh only missing providers
Discovery->>Cache: cache remaining catalogs
```

</a>
<a href="#gh-dark-mode-only">

```mermaid
%%{init: {'theme': 'base', 'themeVariables': {"darkMode": true, "background": "#0d1117", "primaryColor": "#21262d", "primaryTextColor": "#e6edf3", "primaryBorderColor": "#8b949e", "lineColor": "#8b949e", "textColor": "#e6edf3", "edgeLabelBackground": "#161b22", "actorBkg": "#21262d", "actorBorder": "#8b949e", "actorTextColor": "#e6edf3", "actorLineColor": "#8b949e", "signalColor": "#8b949e", "signalTextColor": "#e6edf3", "noteBkgColor": "#373320", "noteBorderColor": "#d4a72c", "noteTextColor": "#f0e6c0", "labelBoxBkgColor": "#21262d", "labelBoxBorderColor": "#8b949e", "labelTextColor": "#e6edf3", "loopTextColor": "#e6edf3", "activationBkgColor": "#30363d", "activationBorderColor": "#8b949e"}}}%%
sequenceDiagram
participant App as ApplicationRuntime
participant Manager as ProviderRuntimeManager
participant Discovery as ProviderModelDiscovery
participant Provider
participant Cache as ProviderModelCache

App->>Manager: warm_referenced_model_cache()
Manager->>Discovery: warm referenced providers
par Provider queries
    Discovery->>Provider: list_model_infos()
end
Provider-->>Discovery: metadata or failure
Discovery->>Cache: cache successful results
Discovery-->>Manager: refresh result
Manager-->>App: warm complete
App->>Manager: start_model_list_refresh()
Manager->>Discovery: refresh only missing providers
Discovery->>Cache: cache remaining catalogs
```

</a>
</details>

<sub>Reviews (1): Last reviewed commit: ["Unify provider model discovery
ownership"](https://github.com/alishahryar1/free-claude-code/commit/33f68e4338fd2326a7d8fd3e7e28246b47fa4210)
| [Re-trigger
Greptile](https://app.greptile.com/api/retrigger?id=45862060)</sub>

<!-- /greptile_comment -->
2026-07-20 23:53:55 -07:00
Ali Khokhar 36bb282558 Recover Claude sessions from provider context overflow (#1215)
## Problem
NVIDIA NIM can report context exhaustion as a 400 `BadRequestError`
saying its derived `max_tokens` is negative. FCC treated that as a
generic invalid request, so Claude Code could not recognize the smaller
upstream context window, compact the conversation, and replay the
interrupted turn. LM Studio also encoded Claude's recovery phrase inside
provider code instead of reporting a protocol-neutral failure. Fixes
#1198.

## Changes
- Add one protocol-neutral `context_window_exceeded` execution failure
with a non-retryable 400 contract.
- Narrowly classify only NVIDIA NIM's negative derived-`max_tokens`
signature while preserving the complete redacted provider diagnostic and
request ID.
- Let the Anthropic serializer alone add Claude's `prompt is too long`
compaction trigger; OpenAI Responses keeps a standard invalid-request
envelope.
- Migrate LM Studio's existing context preflight to the same neutral
semantic and document the ownership boundary.
- Add provider, protocol, API, trace, near-miss, and real Claude
compaction/replay coverage; bump FCC to 4.11.4.

| Before | After |
| --- | --- |
| Context overflow appeared as an ordinary provider 400 and ended the
Claude turn. | Claude receives a typed 400 with its recognized
compaction trigger, compacts once, and replays the interrupted turn
without an FCC or SDK retry loop. |

<!-- greptile_comment -->

<details open><summary><h3>Greptile Summary</h3></summary>

This PR adds protocol-neutral recovery from provider context-window
exhaustion. The main changes are:

- Classify NVIDIA NIM negative derived-`max_tokens` errors as
context-window failures.
- Move Claude's compaction trigger into the Anthropic serializer.
- Migrate LM Studio context preflight to the neutral failure type.
- Preserve the standard OpenAI Responses invalid-request envelope.
- Add provider, protocol, API, trace, and recovery tests.
- Bump the package version to 4.11.4.
</details>

<h3>Confidence Score: 5/5</h3>

This looks safe to merge.

- No blocking issues found in the changed code.
- The new failure kind is covered by both protocol mappings.
- Provider classification remains narrow and non-retryable.
- The protocol-specific compaction phrase stays at the Anthropic
boundary.

<details><summary><h3><a href="https://www.greptile.com/trex"><img
alt="T-Rex"
src="https://greptile-static-assets.s3.amazonaws.com/trex/trex_green.svg"
height="20" align="absmiddle"></a> T-Rex Logs</h3></summary>

**What T-Rex did**
- Ran the narrow pytest command from /home/user/repo without live
provider credentials or services, and observed a clean test run with all
tests passing.

<a
href="https://app.greptile.com/trex/runs/15073403/artifacts"><picture><source
media="(prefers-color-scheme: dark)"
srcset="https://greptile-static-assets.s3.amazonaws.com/badges/ViewAllArtifactsDark.svg?v=4"><source
media="(prefers-color-scheme: light)"
srcset="https://greptile-static-assets.s3.amazonaws.com/badges/ViewAllArtifacts.svg?v=4"><img
alt="View all artifacts"
src="https://greptile-static-assets.s3.amazonaws.com/badges/ViewAllArtifacts.svg?v=4"></picture></a>

<sub><a href="https://www.greptile.com/trex"><img alt="T-Rex"
src="https://greptile-static-assets.s3.amazonaws.com/trex/trex_green.svg"
height="14" align="absmiddle"></a> Ran code and verified through
T-Rex</sub>
</details>

<details open><summary><h3>Important Files Changed</h3></summary>

| Filename | Overview |
|----------|----------|
| src/free_claude_code/providers/nvidia_nim/client.py | Adds narrow
context-window classification for nested and top-level NVIDIA NIM error
bodies. |
| src/free_claude_code/providers/lmstudio/client.py | Migrates
context-budget preflight failures to the canonical context-window
failure. |
| src/free_claude_code/core/anthropic/errors.py | Adds Anthropic mapping
and injects Claude's compaction phrase at the wire boundary. |
| src/free_claude_code/core/openai_responses/errors.py | Maps context
exhaustion to the standard OpenAI invalid-request error. |
| src/free_claude_code/providers/failure_policy.py | Adds the canonical
non-retryable status-400 context-window failure factory. |
| src/free_claude_code/core/failures.py | Adds the protocol-neutral
context-window failure category. |

</details>

<details open><summary><h3>Flowchart</h3></summary>

<a href="#gh-light-mode-only">

```mermaid
%%{init: {'theme': 'neutral'}}%%
flowchart TD
A[Provider detects context exhaustion] --> B[Context-window ExecutionFailure]
B --> C{Protocol adapter}
C -->|Anthropic Messages| D[400 invalid_request_error]
D --> E[Add prompt is too long trigger]
E --> F[Claude compacts and replays]
C -->|OpenAI Responses| G[400 invalid_request_error]
G --> H[Keep neutral provider message]
```

</a>
<a href="#gh-dark-mode-only">

```mermaid
%%{init: {'theme': 'base', 'themeVariables': {"darkMode": true, "background": "#0d1117", "primaryColor": "#21262d", "primaryTextColor": "#e6edf3", "primaryBorderColor": "#8b949e", "lineColor": "#8b949e", "textColor": "#e6edf3", "edgeLabelBackground": "#161b22", "actorBkg": "#21262d", "actorBorder": "#8b949e", "actorTextColor": "#e6edf3", "actorLineColor": "#8b949e", "signalColor": "#8b949e", "signalTextColor": "#e6edf3", "noteBkgColor": "#373320", "noteBorderColor": "#d4a72c", "noteTextColor": "#f0e6c0", "labelBoxBkgColor": "#21262d", "labelBoxBorderColor": "#8b949e", "labelTextColor": "#e6edf3", "loopTextColor": "#e6edf3", "activationBkgColor": "#30363d", "activationBorderColor": "#8b949e"}}}%%
flowchart TD
A[Provider detects context exhaustion] --> B[Context-window ExecutionFailure]
B --> C{Protocol adapter}
C -->|Anthropic Messages| D[400 invalid_request_error]
D --> E[Add prompt is too long trigger]
E --> F[Claude compacts and replays]
C -->|OpenAI Responses| G[400 invalid_request_error]
G --> H[Keep neutral provider message]
```

</a>
</details>

<sub>Reviews (1): Last reviewed commit: ["Normalize provider context
overflow for
..."](https://github.com/alishahryar1/free-claude-code/commit/29c89a4c1011d986284e9d163855174d1d433293)
| [Re-trigger
Greptile](https://app.greptile.com/api/retrigger?id=45603631)</sub>

<!-- /greptile_comment -->
2026-07-20 04:23:28 -07:00
Ali Khokhar fe40548a30 Give Google reasoning controls one owner (#1211)
## Problem

Gemini requests with FCC reasoning could send both `reasoning_effort`
and `extra_body.google.thinking_config`, which [Google documents as
mutually
exclusive](https://ai.google.dev/gemini-api/docs/openai#thinking).
Google reasoning and thought-signature postprocessing had overlapping
request ownership. Fixes #1206.

## Changes

| Before | After |
| --- | --- |
| Shared Google quirks injected thought output independently of the
profile encoder. | One provider-selected Google encoder owns every
reasoning wire field. |
| Gemini could send named effort beside a custom thinking config. |
Gemini selects one channel, with exact budgets taking precedence over
named effort. |
| Vertex reasoning and thought-signature behavior shared one quirks
module. | Vertex retains its budget mapping while thought signatures
have a separate owner. |
| Caller-native Google controls could collide with FCC controls. |
Native controls are preserved only under provider-default reasoning;
controlled collisions fail preflight. |
| Regression coverage inspected isolated body fragments. | Policy
matrices, SDK-merge assertions, and an ownership contract enforce the
final wire shape. |
| The reported Gemini 3.5 Flash path failed upstream with HTTP 400. |
The same live path streams to a normal terminal stop with one reasoning
channel. |

<!-- greptile_comment -->

<details open><summary><h3>Greptile Summary</h3></summary>

This PR gives each Google provider one owner for reasoning request
fields. The main changes are:

- Adds dedicated Gemini and Vertex reasoning encoders.
- Separates thought-signature replay from reasoning serialization.
- Validates caller-provided Google configuration before encoding.
- Adds request-policy and final wire-shape tests.
- Bumps the package version to 4.11.2.
</details>

<h3>Confidence Score: 5/5</h3>

This looks safe to merge.

No blocking issues found in the changed code. The request pipeline keeps
thought signatures and reasoning fields separate. Tests cover the
supported reasoning policies and caller configuration conflicts.

<details><summary><h3><a href="https://www.greptile.com/trex"><img
alt="T-Rex"
src="https://greptile-static-assets.s3.amazonaws.com/trex/trex_green.svg"
height="20" align="absmiddle"></a> T-Rex Logs</h3></summary>

**What T-Rex did**
- The focused validation suite was executed and completed with 78 passed
in 4.57s.
- Before the change, adaptive/high emitted two channels:
reasoning\_effort: high and thinking\_config.include\_thoughts: true.
- After the change, the same request emits exactly one channel:
reasoning\_effort: high, with thinking\_config: null.

<a
href="https://app.greptile.com/trex/runs/15056242/artifacts"><picture><source
media="(prefers-color-scheme: dark)"
srcset="https://greptile-static-assets.s3.amazonaws.com/badges/ViewAllArtifactsDark.svg?v=4"><source
media="(prefers-color-scheme: light)"
srcset="https://greptile-static-assets.s3.amazonaws.com/badges/ViewAllArtifacts.svg?v=4"><img
alt="View all artifacts"
src="https://greptile-static-assets.s3.amazonaws.com/badges/ViewAllArtifacts.svg?v=4"></picture></a>

<sub><a href="https://www.greptile.com/trex"><img alt="T-Rex"
src="https://greptile-static-assets.s3.amazonaws.com/trex/trex_green.svg"
height="14" align="absmiddle"></a> Ran code and verified through
T-Rex</sub>
</details>

<details open><summary><h3>Important Files Changed</h3></summary>

| Filename | Overview |
|----------|----------|
| src/free_claude_code/providers/google_openai/reasoning.py | Adds
exclusive Gemini and Vertex reasoning encoders and validates
caller-native thinking configuration. |
| src/free_claude_code/providers/google_openai/provider.py | Separates
message signature replay from profile-owned reasoning encoding. |
| src/free_claude_code/providers/google_openai/thought_signatures.py |
Narrows the former quirks module to tool-call thought-signature replay.
|
| src/free_claude_code/providers/gemini/client.py | Selects the Gemini
encoder and enables validated extra-body forwarding. |
| src/free_claude_code/providers/vertex/client.py | Selects the Vertex
encoder while retaining budget-based reasoning controls. |
| tests/providers/test_gemini.py | Covers channel exclusivity, budget
precedence, native configuration, conflicts, and SDK merging. |
| tests/providers/test_vertex.py | Covers Vertex policy mapping, native
configuration, conflicts, and final wire shape. |
| tests/contracts/test_import_boundaries.py | Enforces one source owner
for Google reasoning wire fields. |

</details>

<details open><summary><h3>Flowchart</h3></summary>

<a href="#gh-light-mode-only">

```mermaid
%%{init: {'theme': 'neutral'}}%%
flowchart TD
A[Messages request] --> B[Resolve reasoning policy]
A --> C[Validate and copy extra_body]
B --> D{Google provider profile}
C --> E[Build OpenAI request body]
E --> F[Replay thought signatures]
F --> D
D -->|Gemini| G[Gemini reasoning encoder]
D -->|Vertex| H[Vertex reasoning encoder]
G --> I[One Gemini reasoning channel]
H --> J[Google thinking configuration]
I --> K[Final SDK request]
J --> K
```

</a>
<a href="#gh-dark-mode-only">

```mermaid
%%{init: {'theme': 'base', 'themeVariables': {"darkMode": true, "background": "#0d1117", "primaryColor": "#21262d", "primaryTextColor": "#e6edf3", "primaryBorderColor": "#8b949e", "lineColor": "#8b949e", "textColor": "#e6edf3", "edgeLabelBackground": "#161b22", "actorBkg": "#21262d", "actorBorder": "#8b949e", "actorTextColor": "#e6edf3", "actorLineColor": "#8b949e", "signalColor": "#8b949e", "signalTextColor": "#e6edf3", "noteBkgColor": "#373320", "noteBorderColor": "#d4a72c", "noteTextColor": "#f0e6c0", "labelBoxBkgColor": "#21262d", "labelBoxBorderColor": "#8b949e", "labelTextColor": "#e6edf3", "loopTextColor": "#e6edf3", "activationBkgColor": "#30363d", "activationBorderColor": "#8b949e"}}}%%
flowchart TD
A[Messages request] --> B[Resolve reasoning policy]
A --> C[Validate and copy extra_body]
B --> D{Google provider profile}
C --> E[Build OpenAI request body]
E --> F[Replay thought signatures]
F --> D
D -->|Gemini| G[Gemini reasoning encoder]
D -->|Vertex| H[Vertex reasoning encoder]
G --> I[One Gemini reasoning channel]
H --> J[Google thinking configuration]
I --> K[Final SDK request]
J --> K
```

</a>
</details>

<sub>Reviews (1): Last reviewed commit: ["Give Google reasoning controls
one
owner"](https://github.com/alishahryar1/free-claude-code/commit/cdfba8273878d5075ac8d6afb702e0f76224cba4)
| [Re-trigger
Greptile](https://app.greptile.com/api/retrigger?id=45569811)</sub>

<!-- /greptile_comment -->
2026-07-20 00:26:00 -07:00
Ali Khokhar af12e7b2bb Coordinate provider recovery under concurrent load (#1205)
## Problem

Concurrent transient failures could start independent retry, replay,
continuation, and repair loops while holding provider concurrency slots.
This multiplied upstream attempts and could delay or strand terminal
errors under fan-out.

## Changes

| Before | After |
| --- | --- |
| Retry paths owned separate attempt budgets. | One logical-execution
session caps all upstream work at five attempts. |
| Concurrent failures backed off independently. | One provider-owned
recovery episode elects a single half-open probe while followers
coalesce. |
| Backoff occupied stream concurrency. | Concurrency is held only while
an upstream operation or stream is active. |
| Provider catalog calls and stream creation used separate admission
paths. | Every upstream operation uses one provider-generation admission
controller. |
| Cancellation could leave recovery ownership or follower state
unresolved. | Cancellation releases permits, transfers probe ownership,
and unregisters waiting followers. |
| Late in-flight failures could cross an exhausted episode boundary. |
Every coalesced execution retains that generation's terminal outcome. |
| Replay tests allowed loose lifecycle assertions. | Exact SSE contracts
prove retries and continuations emit one unduplicated response. |
| Recovery wrappers could mask final diagnostics. | Final responses and
traces retain the raw provider failure and request ID. |

<!-- greptile_comment -->

<details open><summary><h3>Greptile Summary</h3></summary>

This PR coordinates provider recovery and retry work under concurrent
load. The main changes are:

- One five-attempt budget for each logical execution.
- Provider-wide recovery episodes with one elected probe.
- Shared admission for streams, catalog calls, rate limits, and
concurrency.
- Concurrency permits held only during active upstream work.
- Cancellation-safe probe ownership and preserved final diagnostics.
</details>

<h3>Confidence Score: 5/5</h3>

This looks safe to merge.

No blocking issues found in the changed code.

<details><summary><h3><a href="https://www.greptile.com/trex"><img
alt="T-Rex"
src="https://greptile-static-assets.s3.amazonaws.com/trex/trex_green.svg"
height="20" align="absmiddle"></a> T-Rex Logs</h3></summary>

**What T-Rex did**
- Reviewed the coordinated-recovery-01-before.log to understand how the
exhausted generation outcome was not preserved in a late in-flight
failure.
- Reviewed the coordinated-recovery-02-after.log to confirm that the
updated implementation preserves the exhausted generation outcome for
the same focused contract set.
- Validated that the provider-admission-full-current.log shows the
complete requested test file passed under Python 3.14 with uv run pytest
-n 0.

<a
href="https://app.greptile.com/trex/runs/15050270/artifacts"><picture><source
media="(prefers-color-scheme: dark)"
srcset="https://greptile-static-assets.s3.amazonaws.com/badges/ViewAllArtifactsDark.svg?v=4"><source
media="(prefers-color-scheme: light)"
srcset="https://greptile-static-assets.s3.amazonaws.com/badges/ViewAllArtifacts.svg?v=4"><img
alt="View all artifacts"
src="https://greptile-static-assets.s3.amazonaws.com/badges/ViewAllArtifacts.svg?v=4"></picture></a>

<sub><a href="https://www.greptile.com/trex"><img alt="T-Rex"
src="https://greptile-static-assets.s3.amazonaws.com/trex/trex_green.svg"
height="14" align="absmiddle"></a> Ran code and verified through
T-Rex</sub>
</details>

<details open><summary><h3>Important Files Changed</h3></summary>

| Filename | Overview |
|----------|----------|
| src/free_claude_code/providers/admission.py | Adds shared admission,
retry budgets, recovery episodes, probe election, and cancellation
handling. |
| src/free_claude_code/providers/openai_chat/provider.py | Moves stream
creation, replay, continuation, and repair onto one admission-owned
retry session. |
| src/free_claude_code/providers/stream_recovery.py | Selects replay,
continuation, repair, or final failure using the remaining shared
attempt budget. |
| src/free_claude_code/providers/failure_policy.py | Adds recovery
exhaustion handling and preserves the underlying provider error for
final classification. |
| src/free_claude_code/providers/runtime/factory.py | Creates one
admission controller per provider generation and passes it through
provider factories. |

</details>

<details open><summary><h3>Sequence Diagram</h3></summary>

<a href="#gh-light-mode-only">

```mermaid
%%{init: {'theme': 'neutral'}}%%
sequenceDiagram
participant E as Execution
participant A as Admission controller
participant P as Provider
participant F as Concurrent follower

E->>A: Open attempt
A->>P: Send upstream request
P-->>E: Retryable failure
E->>A: Open recovery episode
F->>A: Request admission
A-->>F: Coalesce and wait
E->>A: Claim probe
A->>P: Send half-open probe
alt Probe succeeds
    P-->>E: Valid response
    E->>A: Close recovery episode
    A-->>F: Release waiter
else Probe fails
    P-->>E: Retryable failure
    E->>A: Schedule next probe or finalize error
end
```

</a>
<a href="#gh-dark-mode-only">

```mermaid
%%{init: {'theme': 'base', 'themeVariables': {"darkMode": true, "background": "#0d1117", "primaryColor": "#21262d", "primaryTextColor": "#e6edf3", "primaryBorderColor": "#8b949e", "lineColor": "#8b949e", "textColor": "#e6edf3", "edgeLabelBackground": "#161b22", "actorBkg": "#21262d", "actorBorder": "#8b949e", "actorTextColor": "#e6edf3", "actorLineColor": "#8b949e", "signalColor": "#8b949e", "signalTextColor": "#e6edf3", "noteBkgColor": "#373320", "noteBorderColor": "#d4a72c", "noteTextColor": "#f0e6c0", "labelBoxBkgColor": "#21262d", "labelBoxBorderColor": "#8b949e", "labelTextColor": "#e6edf3", "loopTextColor": "#e6edf3", "activationBkgColor": "#30363d", "activationBorderColor": "#8b949e"}}}%%
sequenceDiagram
participant E as Execution
participant A as Admission controller
participant P as Provider
participant F as Concurrent follower

E->>A: Open attempt
A->>P: Send upstream request
P-->>E: Retryable failure
E->>A: Open recovery episode
F->>A: Request admission
A-->>F: Coalesce and wait
E->>A: Claim probe
A->>P: Send half-open probe
alt Probe succeeds
    P-->>E: Valid response
    E->>A: Close recovery episode
    A-->>F: Release waiter
else Probe fails
    P-->>E: Retryable failure
    E->>A: Schedule next probe or finalize error
end
```

</a>
</details>

<sub>Reviews (2): Last reviewed commit: ["Harden coordinated retry
lifecycle
invar..."](https://github.com/alishahryar1/free-claude-code/commit/2e871c8649d148b5eb71d21f80bf870ae2d11708)
| [Re-trigger
Greptile](https://app.greptile.com/api/retrigger?id=45554917)</sub>

<!-- /greptile_comment -->
2026-07-19 22:43:29 -07:00
Ali Khokhar a0f62c598c Add Google Vertex AI with renewable ADC (#1193)
## Problem

| Before | After |
| --- | --- |
| FCC supported Google AI Studio API keys but could not route coding
agents through a Google Cloud Vertex AI project. | `vertex/...` routes
through Google's [documented OpenAI-compatible Chat Completions
endpoint](https://cloud.google.com/vertex-ai/generative-ai/docs/start/openai),
using the global endpoint by default or an explicitly configured region.
|
| A pasted Vertex access token would expire, while Application Default
Credentials were not part of provider construction. | FCC loads
[Application Default
Credentials](https://cloud.google.com/docs/authentication/application-default-credentials),
supplies a renewable credential callback to the OpenAI transport,
coalesces concurrent refreshes, and returns typed authentication or
transient failures. |
| Vertex does not expose its model catalog through the compatible OpenAI
`/models` route. | FCC translates its generic discovery operation to
Google's paginated [publisher-model list
API](https://cloud.google.com/vertex-ai/docs/reference/rest/v1beta1/publishers.models/list)
and converts resource names into the model IDs accepted by Chat
Completions. |
| Google thought signatures were owned by the AI Studio adapter even
though Vertex shares the same protocol behavior. | A neutral Google
OpenAI family owns shared thought-signature and request behavior; AI
Studio and Vertex retain separate endpoint and authentication ownership.
|

## Changes

- Added the Vertex provider, `VERTEX_PROJECT_ID`, optional
`VERTEX_LOCATION` and `VERTEX_PROXY`, Admin UI configuration,
model-picker discovery, smoke metadata, and customer setup
documentation.
- Added renewable ADC access tokens with refresh coalescing, proxy-aware
refresh, sanitized failure classification, and project quota headers.
- Added global/regional endpoint composition plus native model-catalog
pagination, strict response validation, response cleanup, and
repeated-page protection.
- Generalized provider readiness around declared configuration fields so
project-based and multi-field providers no longer pretend every remote
provider is configured by one API key.
- Moved shared Google request quirks out of the Gemini adapter,
preserved AI Studio behavior, and bumped the package to `4.11.0`.

<!-- greptile_comment -->

<details open><summary><h3>Greptile Summary</h3></summary>

This PR adds Google Vertex AI as a new provider using Application
Default Credentials. The main changes are:

- New `vertex` provider with project/location endpoint construction.
- Renewable ADC access-token loading with refresh coalescing and
proxy-aware refresh.
- Native Vertex publisher-model discovery with pagination and response
validation.
- Shared Google OpenAI-compatible request behavior for Gemini and
Vertex.
- Admin UI, settings, smoke config, docs, version, lockfile, and tests
for the new provider.
</details>

<h3>Confidence Score: 5/5</h3>

Safe to merge with low risk.

No blocking correctness or security issues were identified. The new
provider follows the existing provider-runtime and Admin configuration
patterns. Endpoint, auth, model parsing, readiness, docs, version,
lockfile, and tests are updated together.

No files require special attention.

<details><summary><h3><a href="https://www.greptile.com/trex"><img
alt="T-Rex"
src="https://greptile-static-assets.s3.amazonaws.com/trex/trex_green.svg"
height="20" align="absmiddle"></a> T-Rex Logs</h3></summary>

**What T-Rex did**
- The T-Rex test suite was executed to validate the code-execution
proof-of-work, generating a full verbose pytest log and recording the
run metadata, and the run completed with EXIT\_CODE: 0.

<a
href="https://app.greptile.com/trex/runs/14991235/artifacts"><picture><source
media="(prefers-color-scheme: dark)"
srcset="https://greptile-static-assets.s3.amazonaws.com/badges/ViewAllArtifactsDark.svg?v=4"><source
media="(prefers-color-scheme: light)"
srcset="https://greptile-static-assets.s3.amazonaws.com/badges/ViewAllArtifacts.svg?v=4"><img
alt="View all artifacts"
src="https://greptile-static-assets.s3.amazonaws.com/badges/ViewAllArtifacts.svg?v=4"></picture></a>

<sub><a href="https://www.greptile.com/trex"><img alt="T-Rex"
src="https://greptile-static-assets.s3.amazonaws.com/trex/trex_green.svg"
height="14" align="absmiddle"></a> Ran code and verified through
T-Rex</sub>
</details>

<details open><summary><h3>Important Files Changed</h3></summary>

| Filename | Overview |
|----------|----------|
| src/free_claude_code/providers/vertex/client.py | Adds the Vertex
provider with OpenAI-compatible chat routing and native paginated model
discovery. |
| src/free_claude_code/providers/vertex/auth.py | Implements renewable
ADC token loading, proxy-aware refresh, coalescing, and sanitized auth
failures. |
| src/free_claude_code/providers/vertex/endpoint.py | Builds validated
Vertex global/regional service, chat, and model-list endpoints. |
| src/free_claude_code/providers/vertex/models.py | Parses Vertex
publisher-model pages into OpenAI-compatible model IDs with
malformed-response checks. |
| src/free_claude_code/providers/google_openai/provider.py | Adds shared
Google thought-signature caching and thinking-budget request body
handling. |
| src/free_claude_code/providers/google_openai/quirks.py | Renames
Gemini-specific quirks to shared Google quirks and exposes model-neutral
thinking config helpers. |
| src/free_claude_code/providers/openai_chat/provider.py | Allows
OpenAI-chat providers to pass an async API-key callback into the OpenAI
SDK. |
| src/free_claude_code/providers/runtime/discovery.py | Uses
descriptor-defined readiness to choose providers eligible for model
cache/discovery. |
| src/free_claude_code/config/provider_catalog.py | Adds the Vertex
descriptor and required settings metadata, and makes Cloudflare
readiness require both token and account ID. |
| src/free_claude_code/config/admin/status.py | Generalizes Admin
provider readiness status to use each descriptor's configuration
attributes. |
| src/free_claude_code/config/admin/provider_manifest.py | Adds Admin UI
fields for Vertex project and location alongside generated provider
fields. |
| tests/providers/test_vertex.py | Adds targeted tests for Vertex
endpoints, ADC token refresh, reasoning mapping, and model discovery
pagination. |

</details>

<details open><summary><h3>Sequence Diagram</h3></summary>

<a href="#gh-light-mode-only">

```mermaid
%%{init: {'theme': 'neutral'}}%%
sequenceDiagram
participant User as User / Admin UI
participant Settings as Settings + Provider Catalog
participant Runtime as Provider Runtime
participant Vertex as VertexProvider
participant ADC as Google ADC
participant OpenAI as OpenAI-compatible Chat Endpoint
participant Models as Vertex Publisher Models API

User->>Settings: Set VERTEX_PROJECT_ID / VERTEX_LOCATION / VERTEX_PROXY
Settings->>Runtime: Descriptor reports vertex configured by project id
Runtime->>Vertex: Construct with project, location, proxy, rate limiter
Vertex->>ADC: Load/refresh Application Default Credentials
ADC-->>Vertex: Renewable access token
Vertex->>OpenAI: Stream chat completion with bearer token + x-goog-user-project
OpenAI-->>Vertex: Streaming chat chunks
Vertex-->>Runtime: Normalized provider stream
Runtime->>Vertex: Refresh model list
Vertex->>Models: GET paginated publishers/google/models
Models-->>Vertex: publisherModels + nextPageToken
Vertex-->>Runtime: Prefixed model IDs for cache/model picker
```

</a>
<a href="#gh-dark-mode-only">

```mermaid
%%{init: {'theme': 'base', 'themeVariables': {"darkMode": true, "background": "#0d1117", "primaryColor": "#21262d", "primaryTextColor": "#e6edf3", "primaryBorderColor": "#8b949e", "lineColor": "#8b949e", "textColor": "#e6edf3", "edgeLabelBackground": "#161b22", "actorBkg": "#21262d", "actorBorder": "#8b949e", "actorTextColor": "#e6edf3", "actorLineColor": "#8b949e", "signalColor": "#8b949e", "signalTextColor": "#e6edf3", "noteBkgColor": "#373320", "noteBorderColor": "#d4a72c", "noteTextColor": "#f0e6c0", "labelBoxBkgColor": "#21262d", "labelBoxBorderColor": "#8b949e", "labelTextColor": "#e6edf3", "loopTextColor": "#e6edf3", "activationBkgColor": "#30363d", "activationBorderColor": "#8b949e"}}}%%
sequenceDiagram
participant User as User / Admin UI
participant Settings as Settings + Provider Catalog
participant Runtime as Provider Runtime
participant Vertex as VertexProvider
participant ADC as Google ADC
participant OpenAI as OpenAI-compatible Chat Endpoint
participant Models as Vertex Publisher Models API

User->>Settings: Set VERTEX_PROJECT_ID / VERTEX_LOCATION / VERTEX_PROXY
Settings->>Runtime: Descriptor reports vertex configured by project id
Runtime->>Vertex: Construct with project, location, proxy, rate limiter
Vertex->>ADC: Load/refresh Application Default Credentials
ADC-->>Vertex: Renewable access token
Vertex->>OpenAI: Stream chat completion with bearer token + x-goog-user-project
OpenAI-->>Vertex: Streaming chat chunks
Vertex-->>Runtime: Normalized provider stream
Runtime->>Vertex: Refresh model list
Vertex->>Models: GET paginated publishers/google/models
Models-->>Vertex: publisherModels + nextPageToken
Vertex-->>Runtime: Prefixed model IDs for cache/model picker
```

</a>
</details>

<sub>Reviews (1): Last reviewed commit: ["feat: add Google Vertex AI
provider"](https://github.com/alishahryar1/free-claude-code/commit/97e753f0772e60377865876ca59b2fd8888d922e)
| [Re-trigger
Greptile](https://app.greptile.com/api/retrigger?id=45405432)</sub>

<!-- /greptile_comment -->
2026-07-18 21:45:07 -07:00
Ali Khokhar af658287bd Add Amazon Bedrock Mantle support (#1192)
## Problem

FCC cannot route coding-agent requests through Amazon Bedrock even
though Bedrock Mantle exposes an OpenAI-compatible streaming Chat
Completions API. Users currently need a separate compatibility layer,
and FCC has no catalog, Admin UI, model-discovery, or smoke-test
contract for Bedrock. Fixes #863.

## Changes

| Before | After |
| --- | --- |
| Amazon Bedrock was absent from provider routing. | `bedrock/` uses the
existing OpenAI Chat provider with AWS's [Bedrock Mantle
endpoint](https://docs.aws.amazon.com/bedrock/latest/userguide/inference-chat-completions-mantle.html).
|
| Bedrock integration would have implied a new AWS-native transport. |
The ordinary profile owns streaming, tools, retries, and `/models`
discovery without boto3, SigV4, Converse, or Invoke machinery. |
| FCC had no Bedrock authentication or regional endpoint configuration.
| `AWS_BEARER_TOKEN_BEDROCK`, `BEDROCK_BASE_URL`, and `BEDROCK_PROXY`
are available through environment and Admin UI configuration, with the
current `us-east-1` Mantle URL as the default. |
| Heterogeneous Bedrock models had no safe provider-wide reasoning
control. | FCC replays prior reasoning through portable think tags and
leaves model-specific reasoning parameters upstream-owned. |
| Smoke configuration duplicated credential checks for every provider. |
Smoke configuration reads primary credentials and configurable endpoints
from the provider catalog, retaining only Cloudflare's two-field
exception. |
| Bedrock behavior had no deterministic coverage or release
documentation. | Provider requests, regional URL normalization, model
discovery, Admin persistence, smoke selection, README usage,
architecture boundaries, and version 4.10.0 cover the new capability. |

<!-- greptile_comment -->

<details open><summary><h3>Greptile Summary</h3></summary>

This PR adds Amazon Bedrock Mantle as an OpenAI-compatible provider. The
main changes are:

- Bedrock catalog, settings, proxy, and regional base-URL configuration.
- OpenAI Chat routing with portable reasoning replay and model
discovery.
- Admin UI fields and configuration persistence.
- Catalog-driven smoke configuration and Bedrock smoke coverage.
- Provider documentation and a version bump to 4.10.0.
</details>

<h3>Confidence Score: 5/5</h3>

The provider flow looks mergeable after handling an explicitly empty
Bedrock base URL.

Catalog, runtime, Admin, and smoke wiring are consistent. Regional URLs
with or without `/v1` are normalized correctly.

src/free_claude_code/config/settings.py: an empty `BEDROCK_BASE_URL` can
still create a client with an invalid endpoint.

<details><summary><h3><a href="https://www.greptile.com/trex"><img
alt="T-Rex"
src="https://greptile-static-assets.s3.amazonaws.com/trex/trex_green.svg"
height="20" align="absmiddle"></a> T-Rex Logs</h3></summary>

**What T-Rex did**
- Before the change, bedrock was rejected as an unknown provider with
exit code 1.
- After the change, the fake service captured the normalized URL, Bearer
trex...oken, portable Chat Completions JSON, one tool, think-tag
history, and no reasoning\_effort, reasoning, or thinking fields; exit
code 0.
- Focused pytest validation completed with 6/6 passing tests.

<a
href="https://app.greptile.com/trex/runs/14988686/artifacts"><picture><source
media="(prefers-color-scheme: dark)"
srcset="https://greptile-static-assets.s3.amazonaws.com/badges/ViewAllArtifactsDark.svg?v=4"><source
media="(prefers-color-scheme: light)"
srcset="https://greptile-static-assets.s3.amazonaws.com/badges/ViewAllArtifacts.svg?v=4"><img
alt="View all artifacts"
src="https://greptile-static-assets.s3.amazonaws.com/badges/ViewAllArtifacts.svg?v=4"></picture></a>

<sub><a href="https://www.greptile.com/trex"><img alt="T-Rex"
src="https://greptile-static-assets.s3.amazonaws.com/trex/trex_green.svg"
height="14" align="absmiddle"></a> Ran code and verified through
T-Rex</sub>
</details>

<details open><summary><h3>Important Files Changed</h3></summary>

| Filename | Overview |
|----------|----------|
| src/free_claude_code/config/provider_catalog.py | Adds the Bedrock
provider descriptor, regional default endpoint, credential mapping, and
proxy metadata. |
| src/free_claude_code/config/settings.py | Adds Bedrock settings, but
an explicitly empty base URL bypasses the regional default. |
| src/free_claude_code/providers/openai_chat/profiles.py | Registers
Bedrock with URL normalization, think-tag replay, and no provider-wide
reasoning parameter. |
| src/free_claude_code/config/admin/provider_manifest.py | Generalizes
provider base-URL fields and adds Bedrock-specific labels and help text.
|
| smoke/lib/config.py | Adds Bedrock smoke defaults and replaces
provider-specific checks with catalog-driven configuration checks. |
| tests/providers/test_bedrock.py | Covers Bedrock URL normalization,
request fields, reasoning replay, tool calls, and model discovery. |

</details>

<details open><summary><h3>Flowchart</h3></summary>

<a href="#gh-light-mode-only">

```mermaid
%%{init: {'theme': 'neutral'}}%%
flowchart LR
Config[Environment or Admin config] --> Settings[Settings]
Catalog[Provider catalog] --> Runtime[Provider runtime]
Settings --> Runtime
Runtime --> Profile[Bedrock OpenAI Chat profile]
Profile --> Client[AsyncOpenAI client]
Client --> Mantle[Regional Bedrock Mantle endpoint]
Catalog --> Admin[Admin manifest]
Catalog --> Smoke[Smoke selection]
```

</a>
<a href="#gh-dark-mode-only">

```mermaid
%%{init: {'theme': 'base', 'themeVariables': {"darkMode": true, "background": "#0d1117", "primaryColor": "#21262d", "primaryTextColor": "#e6edf3", "primaryBorderColor": "#8b949e", "lineColor": "#8b949e", "textColor": "#e6edf3", "edgeLabelBackground": "#161b22", "actorBkg": "#21262d", "actorBorder": "#8b949e", "actorTextColor": "#e6edf3", "actorLineColor": "#8b949e", "signalColor": "#8b949e", "signalTextColor": "#e6edf3", "noteBkgColor": "#373320", "noteBorderColor": "#d4a72c", "noteTextColor": "#f0e6c0", "labelBoxBkgColor": "#21262d", "labelBoxBorderColor": "#8b949e", "labelTextColor": "#e6edf3", "loopTextColor": "#e6edf3", "activationBkgColor": "#30363d", "activationBorderColor": "#8b949e"}}}%%
flowchart LR
Config[Environment or Admin config] --> Settings[Settings]
Catalog[Provider catalog] --> Runtime[Provider runtime]
Settings --> Runtime
Runtime --> Profile[Bedrock OpenAI Chat profile]
Profile --> Client[AsyncOpenAI client]
Client --> Mantle[Regional Bedrock Mantle endpoint]
Catalog --> Admin[Admin manifest]
Catalog --> Smoke[Smoke selection]
```

</a>
</details>

<a
href="https://app.greptile.com/api/ide/codex?prompt=IMPORTANT%3A%20Work%20in%20the%20repository%20%22alishahryar1%2Ffree-claude-code%22%20on%20the%20existing%20branch%20%22ali%2Fadd-bedrock-mantle%22.%20Checkout%20that%20branch%20%E2%80%94%20do%20NOT%20create%20a%20new%20branch%20or%20open%20a%20new%20PR.%20Push%20your%20changes%20to%20%22ali%2Fadd-bedrock-mantle%22.%0A%0AFix%20the%20following%201%20code%20review%20issue.%20Work%20through%20them%20one%20at%20a%20time%2C%20proposing%20concise%20fixes.%0A%0A---%0A%0A%23%23%23%20Issue%201%20of%201%0Asrc%2Ffree_claude_code%2Fconfig%2Fsettings.py%3A60-63%0A**Empty%20Override%20Bypasses%20Default%20URL**%0A%0AWhen%20%60BEDROCK_BASE_URL%60%20is%20present%20but%20empty%2C%20settings%20keep%20the%20empty%20string%20instead%20of%20using%20%60BEDROCK_DEFAULT_BASE%60.%20Clearing%20this%20field%20through%20environment%20or%20Admin%20configuration%20therefore%20builds%20the%20Bedrock%20client%20with%20an%20invalid%20base%20URL%2C%20and%20Bedrock%20requests%20and%20model%20discovery%20fail%20instead%20of%20using%20the%20documented%20default.%0A%0A&repo=alishahryar1%2Ffree-claude-code&pr=1192&platform=github"><picture><source
media="(prefers-color-scheme: dark)"
srcset="https://greptile-static-assets.s3.amazonaws.com/badges/FixAllInCodexDark.svg?v=6"><source
media="(prefers-color-scheme: light)"
srcset="https://greptile-static-assets.s3.amazonaws.com/badges/FixAllInCodex.svg?v=6"><img
alt="Fix All in Codex"
src="https://greptile-static-assets.s3.amazonaws.com/badges/FixAllInCodex.svg?v=6"></picture></a>

<sub>Reviews (1): Last reviewed commit: ["Add Amazon Bedrock Mantle
provider"](https://github.com/alishahryar1/free-claude-code/commit/f40b539c99c9f6a888762b8da69d5a738a4f4d70)
| [Re-trigger
Greptile](https://app.greptile.com/api/retrigger?id=45400300)</sub>

> Greptile also left **1 inline comment** on this PR.

<!-- /greptile_comment -->
2026-07-18 20:29:41 -07:00
Ali Khokhar b89e849fff Add Kimi Code subscription support (#1183)
## Problem

FCC's existing `kimi` provider targets Kimi Open Platform credits. Kimi
Code subscription keys use a separate coding-agent endpoint, so
subscribers cannot currently use their plan or select its K3 and coding
models. Fixes #1161.

## Changes

| Before | After |
| --- | --- |
| `KIMI_API_KEY` was the only Kimi contract and routed to the
credit-based API platform. | `KIMI_API_KEY` remains unchanged, while
`KIMI_CODE_API_KEY` routes the separate [Kimi Code
subscription](https://www.kimi.com/code/docs/en/) through
`https://api.kimi.com/coding/v1`. |
| OpenAI-chat profiles could not declare an upstream client identity. |
The Kimi Code profile sends an honest `free-claude-code` user agent as
required by [Kimi's integration
policy](https://www.kimi.com/code/docs/en/kimi-code/community-guidelines.html),
without adding a specialized provider class. |
| FCC's fallback output limit and generic token field would override
Kimi's default when the client omitted a limit. | Explicit client limits
become `max_completion_tokens`; omitted limits remain upstream-owned,
and learned cap recovery still handles model-specific rejections. |
| Kimi Code reasoning and history had no provider contract. | The
profile maps FCC intent to Kimi's documented `low`, `high`, `max`, and
`none` efforts and replays prior thinking through `reasoning_content`,
with no model-name branching. |
| Admin, model discovery, and smoke coverage knew only the credit-based
Kimi provider. | Catalog-derived Admin fields, `/models` discovery,
README setup, smoke configuration, and deterministic provider/runtime
contracts cover `kimi_code`; the release is bumped to 4.9.0. |

<!-- greptile_comment -->

<details open><summary><h3>Greptile Summary</h3></summary>

This PR adds Kimi Code subscription support as a separate provider. The
main changes are:

- Adds dedicated API key, proxy, endpoint, and Admin settings.
- Adds Kimi-specific reasoning, token-limit, history, and user-agent
behavior.
- Extends model discovery and smoke configuration for `kimi_code`.
- Adds provider contract tests and updates documentation.
- Bumps the package and lockfile version to 4.9.0.
</details>

<h3>Confidence Score: 5/5</h3>

This looks safe to merge.

No blocking issues found in the changed code.

<details><summary><h3><a href="https://www.greptile.com/trex"><img
alt="T-Rex"
src="https://greptile-static-assets.s3.amazonaws.com/trex/trex_green.svg"
height="20" align="absmiddle"></a> T-Rex Logs</h3></summary>

**What T-Rex did**
- The contract validation compared the Kimi code HTTP flow before and
after the change and confirmed the after state exits with code 0 and
includes Bearer authentication, User-Agent: free-claude-code, model
discovery, max\_completion\_tokens, Kimi reasoning effort max,
reasoning\_content replay, and omission of token/reasoning fields when
unspecified.
- Focused tests for the Kimi code HTTP flow were run, and 88 tests
passed with exit code 0.
- Artifacts document both the initial failure scenario and the
successful post-change state, along with the focused tests results.

<a
href="https://app.greptile.com/trex/runs/14945625/artifacts"><picture><source
media="(prefers-color-scheme: dark)"
srcset="https://greptile-static-assets.s3.amazonaws.com/badges/ViewAllArtifactsDark.svg?v=4"><source
media="(prefers-color-scheme: light)"
srcset="https://greptile-static-assets.s3.amazonaws.com/badges/ViewAllArtifacts.svg?v=4"><img
alt="View all artifacts"
src="https://greptile-static-assets.s3.amazonaws.com/badges/ViewAllArtifacts.svg?v=4"></picture></a>

<sub><a href="https://www.greptile.com/trex"><img alt="T-Rex"
src="https://greptile-static-assets.s3.amazonaws.com/trex/trex_green.svg"
height="14" align="absmiddle"></a> Ran code and verified through
T-Rex</sub>
</details>

<details open><summary><h3>Important Files Changed</h3></summary>

| Filename | Overview |
|----------|----------|
| src/free_claude_code/config/provider_catalog.py | Registers Kimi Code
with its dedicated credential, endpoint, and proxy setting. |
| src/free_claude_code/config/settings.py | Adds environment-backed
fields for the Kimi Code subscription key and proxy. |
| src/free_claude_code/providers/openai_chat/__init__.py | Passes an
optional profile user agent into the shared OpenAI-compatible client. |
| src/free_claude_code/providers/openai_chat/profiles.py | Defines Kimi
Code reasoning, token-limit, history replay, extra-body, and user-agent
behavior. |
| smoke/lib/config.py | Adds Kimi Code configuration detection and its
default smoke model. |
| tests/providers/test_kimi_code.py | Covers the new endpoint, headers,
request mapping, reasoning behavior, and model discovery. |

</details>

<sub>Reviews (1): Last reviewed commit: ["Add Kimi Code subscription
provider"](https://github.com/alishahryar1/free-claude-code/commit/dfc66f2755e5885775a77f5c91b4b9a82eb8fd1d)
| [Re-trigger
Greptile](https://app.greptile.com/api/retrigger?id=45309322)</sub>

<!-- /greptile_comment -->
2026-07-18 04:10:02 -07:00
Ali Khokhar 65a342ede4 Honor reasoning tiers on numeric-budget providers (#1150)
## Problem

FCC preserved named reasoning effort but numeric-budget providers
received no intensity unless a client supplied an exact token budget.
NIM and llama.cpp therefore treated Low through Max like the same
provider default.

## Changes

| Before | After |
| --- | --- |
| Named effort had no FCC-owned numeric meaning. | FCC maps Low, Medium,
High, X-High, and Max to 512, 1,024, 2,048, 4,096, and 8,192 tokens. |
| NIM received only thinking booleans for named effort. |
[NIM](https://docs.nvidia.com/nim/large-language-models/1.15.0/thinking-budget-control.html)
receives thinking booleans plus the mapped `reasoning_budget`, retaining
its existing retry without rejected budget control. |
| llama.cpp forwarded only exact client budgets. |
[llama.cpp](https://github.com/ggml-org/llama.cpp/blob/master/tools/server/README.md)
receives the mapped `thinking_budget_tokens` value. |
| Named, boolean, and provider-default adapters shared no explicit
numeric contract. | Named adapters keep words, boolean adapters keep
on/off, and only numeric-budget adapters consume the FCC scale. |
| Documentation prohibited every named-effort budget conversion. |
Documentation defines the product scale, exact-budget precedence, and
model-independent ownership boundary. |
| Version 4.8.0 exposed tiers that collapsed on numeric providers. |
Version 4.8.1 completes the tiers; all five local CI checks pass with
2,383 tests passed and 40 skipped. |

<!-- greptile_comment -->

<details open><summary><h3>Greptile Summary</h3></summary>

This PR gives named reasoning tiers numeric budgets for providers that
require token counts. The main changes are:

- Adds one shared tier-to-token scale with exact-budget precedence.
- Sends mapped budgets to NVIDIA NIM and llama.cpp.
- Removes conflicting client-supplied NIM budget fields.
- Adds focused provider and policy tests.
- Updates documentation and bumps the package to 4.8.1.
</details>

<h3>Confidence Score: 5/5</h3>

This looks safe to merge.

The NIM fix removes both conflicting budget locations while preserving
unrelated nested options.

No blocking issue remains in the changed paths.

<details><summary><h3><a href="https://www.greptile.com/trex"><img
alt="T-Rex"
src="https://greptile-static-assets.s3.amazonaws.com/trex/trex_green.svg"
height="20" align="absmiddle"></a> T-Rex Logs</h3></summary>

**What T-Rex did**
- T-Rex ran a pre-change focused validation against the budget logic and
observed 16 failures and 2 passes, establishing the baseline.
- T-Rex ran a post-change focused validation (head run) and confirmed
all 18 cases passed, including the mocked NIM retry paths.
- T-Rex executed the full focused validation after the change and
produced a verbose log showing 91/91 passing.

<a
href="https://app.greptile.com/trex/runs/14800553/artifacts"><picture><source
media="(prefers-color-scheme: dark)"
srcset="https://greptile-static-assets.s3.amazonaws.com/badges/ViewAllArtifactsDark.svg?v=4"><source
media="(prefers-color-scheme: light)"
srcset="https://greptile-static-assets.s3.amazonaws.com/badges/ViewAllArtifacts.svg?v=4"><img
alt="View all artifacts"
src="https://greptile-static-assets.s3.amazonaws.com/badges/ViewAllArtifacts.svg?v=4"></picture></a>

<sub><a href="https://www.greptile.com/trex"><img alt="T-Rex"
src="https://greptile-static-assets.s3.amazonaws.com/trex/trex_green.svg"
height="14" align="absmiddle"></a> Ran code and verified through
T-Rex</sub>
</details>

<details open><summary><h3>Important Files Changed</h3></summary>

| Filename | Overview |
|----------|----------|
| src/free_claude_code/core/reasoning.py | Adds the shared
effort-to-token scale and preserves exact client budgets as the
higher-priority value. |
| src/free_claude_code/providers/nvidia_nim/request_options.py | Removes
top-level and nested client budgets before inserting one policy-derived
NIM budget. |
| src/free_claude_code/providers/openai_chat/reasoning.py | Extends
llama.cpp request encoding to send mapped named-effort budgets. |

</details>

<sub>Reviews (2): Last reviewed commit: ["fix: canonicalize NIM
reasoning
budgets"](https://github.com/alishahryar1/free-claude-code/commit/368f1a88d3f1404fcab2365703cca89f1077f1ea)
| [Re-trigger
Greptile](https://app.greptile.com/api/retrigger?id=44998377)</sub>

<!-- /greptile_comment -->
2026-07-16 22:18:20 -07:00
Ali Khokhar 0de7608b52 Keep inline system reminders cache-stable across providers (#1154)
## Problem

Inline Anthropic system reminders were forwarded as mid-conversation
OpenAI system roles. Compatible provider chat templates could reposition
those roles, changing prior prompt tokens and causing periodic full
cache misses. Fixes #1152.

## Changes

| Before | After |
| --- | --- |
| Top-level and inline system content both used downstream system roles.
| Only top-level system content uses the leading downstream system role.
|
| Inline system reminders could trigger provider-side prompt
retemplating. | Inline system reminders keep their text and position as
downstream user messages. |
| Provider policies could reinterpret the shared role mapping. | Shared
conversion owns one provider-independent role mapping. |
| Cache-prefix coverage allowed mid-conversation system roles. |
Cache-prefix coverage requires append-only user-encoded reminders. |
| The package version was 4.8.0. | The package version is 4.8.1. |

<!-- greptile_comment -->

<details open><summary><h3>Greptile Summary</h3></summary>

This PR keeps inline system reminders stable across OpenAI-compatible
providers. The main changes are:

- Maps only the top-level system prompt to the downstream system role.
- Encodes inline system reminders as ordered user content.
- Coalesces adjacent user messages after transcript ordering.
- Adds tests for text, multimodal content, cache-prefix stability, and
tool-result ordering.
- Updates the architecture notes and bumps the package to 4.8.1.
</details>

<h3>Confidence Score: 5/5</h3>

This looks safe to merge.

Adjacent user messages are combined after assistant and tool
dependencies are ordered. Multimodal content keeps its part order. No
blocking issues were found in the changed code.

<details><summary><h3><a href="https://www.greptile.com/trex"><img
alt="T-Rex"
src="https://greptile-static-assets.s3.amazonaws.com/trex/trex_green.svg"
height="20" align="absmiddle"></a> T-Rex Logs</h3></summary>

**What T-Rex did**
- Before capture, the converted prefix roles were shown as system, user,
and assistant.
- After capture, the same prefix appeared, followed by a user
continuation containing a second question and a system-reminder tag; all
runtime assertions passed.
- The preserved harness documents and reproduces the public-builder
validation.

<a
href="https://app.greptile.com/trex/runs/14803290/artifacts"><picture><source
media="(prefers-color-scheme: dark)"
srcset="https://greptile-static-assets.s3.amazonaws.com/badges/ViewAllArtifactsDark.svg?v=4"><source
media="(prefers-color-scheme: light)"
srcset="https://greptile-static-assets.s3.amazonaws.com/badges/ViewAllArtifacts.svg?v=4"><img
alt="View all artifacts"
src="https://greptile-static-assets.s3.amazonaws.com/badges/ViewAllArtifacts.svg?v=4"></picture></a>

<sub><a href="https://www.greptile.com/trex"><img alt="T-Rex"
src="https://greptile-static-assets.s3.amazonaws.com/trex/trex_green.svg"
height="14" align="absmiddle"></a> Ran code and verified through
T-Rex</sub>
</details>

<details open><summary><h3>Important Files Changed</h3></summary>

| Filename | Overview |
|----------|----------|
| src/free_claude_code/core/anthropic/conversion.py | Maps inline system
reminders to user content and coalesces adjacent user messages after
transcript ordering. |
| tests/providers/test_converter.py | Adds tests for inline reminders,
multimodal content, cache-prefix stability, and tool-result ordering. |
| ARCHITECTURE.md | Documents the shared role-mapping and adjacent-user
coalescing rules. |
| pyproject.toml | Bumps the package version to 4.8.1. |
| uv.lock | Synchronizes the locked editable package version with 4.8.1.
|

</details>

<sub>Reviews (2): Last reviewed commit: ["fix: coalesce adjacent
provider user
tur..."](https://github.com/alishahryar1/free-claude-code/commit/7af6327c79fa15c0f4922bad8864f1c6656b2814)
| [Re-trigger
Greptile](https://app.greptile.com/api/retrigger?id=45005803)</sub>

<!-- /greptile_comment -->
2026-07-16 22:14:33 -07:00
Ali Khokhar 6455c63e1d Make reasoning policy provider-neutral and client-aware (#1148)
## Problem

FCC reduced reasoning to global and route booleans, mixing client
intent, configuration, provider wire capabilities, output visibility,
and history replay. That discarded named client efforts, encouraged
model-name checks, and made provider behavior inconsistent.

## Changes

| Before | After |
| --- | --- |
| Admin exposed global and route thinking toggles. | Admin exposes
**Off**, **From client**, **Low**, **Medium**, **High**, **X-High**, and
**Max**; Fable, Opus, Sonnet, and Haiku also expose **Inherit**. |
| Request intent was repeatedly reduced to a boolean across routing and
providers. | The application boundary resolves one immutable
`ReasoningPolicy` with independent control, named effort, and exact
positive token budget. |
| Provider adapters could infer reasoning behavior from upstream model
names or versions. | Provider profiles translate only documented
provider-wide wire capabilities; architecture and contributor rules
prohibit model-specific reasoning branches. |
| Gateway reasoning controls were ad hoc. |
[OpenRouter](https://openrouter.ai/docs/guides/best-practices/reasoning-tokens)
and [Vercel AI
Gateway](https://vercel.com/docs/ai-gateway/models-and-providers) use
documented reasoning objects, including exact budgets where
representable. |
| Named effort forwarding was inconsistent or absent. |
[Gemini](https://ai.google.dev/gemini-api/docs/openai),
[Ollama](https://docs.ollama.com/api/openai-compatibility), [LM
Studio](https://lmstudio.ai/changelog/lmstudio-v0.4.8),
[Fireworks](https://docs.fireworks.ai/guides/querying-text-models/reasoning),
[Cohere](https://docs.cohere.com/docs/compatibility-api),
[Wafer](https://docs.wafer.ai/serverless/api-reference),
[Groq](https://console.groq.com/docs/reasoning),
[Cerebras](https://inference-docs.cerebras.ai/capabilities/reasoning),
[SambaNova](https://docs.sambanova.ai/docs/api-reference/chat-completions/create-chat-based-completion),
and
[Mistral](https://docs.mistral.ai/studio-api/conversations/reasoning)
receive their documented named vocabularies with explicit provider-owned
downgrades. |
| Boolean thinking controls were mixed into shared conversion. |
[DeepSeek](https://api-docs.deepseek.com/guides/thinking_mode/),
[Kimi](https://platform.kimi.ai/docs/guide/use-kimi-k2-thinking-model),
[Z.ai](https://docs.z.ai/guides/capabilities/thinking-mode), [Cloudflare
Workers
AI](https://developers.cloudflare.com/changelog/post/2026-04-20-kimi-k2-6-workers-ai/),
and [NVIDIA
NIM](https://docs.nvidia.com/nim/large-language-models/1.15.0/thinking-budget-control.html)
use provider-owned thinking-object or chat-template controls. |
| Effort names and output limits could become fabricated reasoning
budgets. | Exact budgets remain exact and are forwarded only through
documented fields for OpenRouter, Fireworks, LM Studio, NIM, and
[llama.cpp](https://github.com/ggml-org/llama.cpp/blob/master/tools/server/README.md);
named efforts and output limits are never converted into token budgets.
|
| New-turn reasoning and prior-turn replay shared one switch. | Every
profile independently declares native reasoning replay, `<think>` tag
replay, provider-specific replay, or no replay; **Off** suppresses new
reasoning output without corrupting required history. |
| Providers without a stable generic compute control received guessed
controls. |
[MiniMax](https://platform.minimax.io/docs/api-reference/text-openai-api)
requests split output only, while [GitHub
Models](https://docs.github.com/en/rest/models/inference), [Hugging Face
Inference
Providers](https://huggingface.co/docs/inference-providers/en/tasks/chat-completion),
Codestral, and OpenCode keep provider defaults and use only their
explicit replay profile. |
| OpenAI Responses effort became a lossy Anthropic thinking boolean. |
Responses preserves `reasoning.effort` through `output_config`, then
resolves it through the same application policy as Messages without
inventing a budget. |
| Legacy booleans remained the persisted contract. | FCC-owned dotenv
files migrate to typed `REASONING_*` values, explicit env files receive
an actionable warning, documentation describes the ownership boundary,
and the package advances to 4.8.0. |
| Reasoning behavior was covered by scattered boolean assertions. | New
policy, routing, encoder, provider, Admin, migration, Responses, and
smoke contracts pass all five local CI checks: 2,368 tests passed, 40
skipped; 92 smoke tests collect and both live config migration checks
pass. |

<!-- greptile_comment -->

<details open><summary><h3>Greptile Summary</h3></summary>

This PR makes reasoning policy client-aware and independent of provider
model names. The main changes are:

- Adds one immutable reasoning policy resolved at the application
boundary.
- Adds typed root and route reasoning settings with Admin UI support.
- Moves wire controls and history replay behavior into provider
profiles.
- Migrates owned dotenv files from legacy thinking booleans.
- Expands provider, routing, migration, API, and smoke coverage.
</details>

<h3>Confidence Score: 5/5</h3>

This looks safe to merge.

No blocking issues found in the changed code.

<details><summary><h3><a href="https://www.greptile.com/trex"><img
alt="T-Rex"
src="https://greptile-static-assets.s3.amazonaws.com/trex/trex_green.svg"
height="20" align="absmiddle"></a> T-Rex Logs</h3></summary>

**What T-Rex did**
- Ran the contract-validation test suite with the specified test
modules, and the tests reported 78 passed in 1.53s with exit code 0.
- Reviewed the complete captured output artifact
reasoning-contract-02-after.log to verify the final test outcomes and
successful contract validation.

<a
href="https://app.greptile.com/trex/runs/14792858/artifacts"><picture><source
media="(prefers-color-scheme: dark)"
srcset="https://greptile-static-assets.s3.amazonaws.com/badges/ViewAllArtifactsDark.svg?v=4"><source
media="(prefers-color-scheme: light)"
srcset="https://greptile-static-assets.s3.amazonaws.com/badges/ViewAllArtifacts.svg?v=4"><img
alt="View all artifacts"
src="https://greptile-static-assets.s3.amazonaws.com/badges/ViewAllArtifacts.svg?v=4"></picture></a>

<sub><a href="https://www.greptile.com/trex"><img alt="T-Rex"
src="https://greptile-static-assets.s3.amazonaws.com/trex/trex_green.svg"
height="14" align="absmiddle"></a> Ran code and verified through
T-Rex</sub>
</details>

<details open><summary><h3>Important Files Changed</h3></summary>

| Filename | Overview |
|----------|----------|
| src/free_claude_code/config/env_migrations.py | Migrates legacy
reasoning booleans in owned dotenv files and warns for explicit
environment files. |
| src/free_claude_code/application/reasoning.py | Resolves client
controls and configured preferences into one provider-neutral reasoning
policy. |
| src/free_claude_code/application/routing.py | Carries route-level
reasoning preferences into request-scoped policy resolution. |
| src/free_claude_code/providers/openai_chat/reasoning.py | Provides
shared provider encoders for reasoning controls and replay behavior. |

</details>

<sub>Reviews (2): Last reviewed commit: ["chore: release reasoning
controls as
4.8..."](https://github.com/alishahryar1/free-claude-code/commit/9d4be767f7dbdca5709474012f43dcdc6f4347e3)
| [Re-trigger
Greptile](https://app.greptile.com/api/retrigger?id=44984039)</sub>

<!-- /greptile_comment -->
2026-07-16 19:57:12 -07:00
Ali Khokhar cba7ed23c5 Fix Cerebras reasoning serialization (#1138)
## Problem

Cerebras rejects multi-turn Claude conversations because FCC serializes
prior assistant thinking as the unsupported `reasoning_content` field.
Cerebras also streams reasoning through `delta.reasoning`, so FCC can
miss reasoning output. Fixes #1136.

## Changes

| Before | After |
| --- | --- |
| Cerebras replayed assistant thinking through `reasoning_content`. |
Cerebras replays assistant thinking as `<think>`-tagged assistant
content. |
| Cerebras parsed streamed reasoning from `delta.reasoning_content`. |
Cerebras parses streamed reasoning from `delta.reasoning`. |
| The disabled-thinking test checked an impossible message role. |
Regression tests inspect serialized fields, preserve tool history, and
cover disabled thinking. |
| The package version was `4.7.1`. | The package version is `4.7.2`. |

<!-- greptile_comment -->

<details open><summary><h3>Greptile Summary</h3></summary>

This PR fixes reasoning serialization for Cerebras conversations. The
main changes are:

- Replays prior assistant reasoning as `<think>`-tagged content.
- Reads streamed reasoning from `delta.reasoning`.
- Adds tests for tool history and disabled thinking.
- Bumps the package version to `4.7.2`.
</details>

<h3>Confidence Score: 5/5</h3>

This looks safe to merge.

- No blocking issues found in the changed code.

<details><summary><h3><a href="https://www.greptile.com/trex"><img
alt="T-Rex"
src="https://greptile-static-assets.s3.amazonaws.com/trex/trex_green.svg"
height="20" align="absmiddle"></a> T-Rex Logs</h3></summary>

**What T-Rex did**
- The pre-change run showed HEAD^ failed the tagged-reasoning replay and
delta.reasoning streaming tests, with 2 failures and 9 successes.
- The post-change head checkout completed successfully, with 11 tests
passing in 0.73s and exit code 0.

<a
href="https://app.greptile.com/trex/runs/14672360/artifacts"><picture><source
media="(prefers-color-scheme: dark)"
srcset="https://greptile-static-assets.s3.amazonaws.com/badges/ViewAllArtifactsDark.svg?v=4"><source
media="(prefers-color-scheme: light)"
srcset="https://greptile-static-assets.s3.amazonaws.com/badges/ViewAllArtifacts.svg?v=4"><img
alt="View all artifacts"
src="https://greptile-static-assets.s3.amazonaws.com/badges/ViewAllArtifacts.svg?v=4"></picture></a>

<sub><a href="https://www.greptile.com/trex"><img alt="T-Rex"
src="https://greptile-static-assets.s3.amazonaws.com/trex/trex_green.svg"
height="14" align="absmiddle"></a> Ran code and verified through
T-Rex</sub>
</details>

<details open><summary><h3>Important Files Changed</h3></summary>

| Filename | Overview |
|----------|----------|
| src/free_claude_code/providers/openai_chat/profiles.py | Updates
Cerebras reasoning replay and streaming field configuration. |
| tests/providers/test_cerebras.py | Covers tagged reasoning replay,
tool history, disabled thinking, and streamed reasoning. |
| pyproject.toml | Bumps the package patch version to 4.7.2. |
| uv.lock | Synchronizes the locked editable package version. |

</details>

<sub>Reviews (1): Last reviewed commit: ["Fix Cerebras reasoning
serialization"](https://github.com/alishahryar1/free-claude-code/commit/bff27ef4de4c968408e6a049bc6cca2068c7c836)
| [Re-trigger
Greptile](https://app.greptile.com/api/retrigger?id=44744399)</sub>

<!-- /greptile_comment -->
2026-07-16 10:22:20 -07:00
Ali Khokhar 8d7f560589 Preserve DeepSeek cache prefixes across tool turns (#1126)
## Problem

DeepSeek cache accounting was restored, but a later tool turn could
still make FCC rewrite older assistant history. That prevented FCC from
guaranteeing the identical serialized prefix required for cache reuse
and left the remaining behavior in #904 unresolved.

## Changes

| Before | After |
| --- | --- |
| Non-tool reasoning was replayed until a later tool fallback removed
it. | Non-tool reasoning is omitted consistently from its first history
serialization. |
| Current-generation thinking controlled all historical reasoning
replay. | Required tool-call reasoning is replayed independently of
current-generation thinking. |
| One replayable tool turn made the entire tool history appear safe. |
Every tool-call turn must have replayable reasoning before thinking
remains enabled. |
| Prefix behavior was not verified after SDK request serialization. |
DeepSeek tests capture final SDK JSON and require an exact append-only
message prefix. |

<!-- greptile_comment -->

<details open><summary><h3>Greptile Summary</h3></summary>

This PR preserves DeepSeek cache prefixes while retaining required
tool-call reasoning. The main changes are:

- Omits non-tool reasoning from its first history serialization.
- Replays tool-call reasoning independently of current thinking mode.
- Disables thinking unless every historical tool call has replayable
reasoning.
- Tests the final SDK-serialized message prefix.
- Bumps the package patch version and updates the lockfile.
</details>

<h3>Confidence Score: 4/5</h3>

The reasoning-only assistant path needs a fix before merging.

The main prefix-preservation flow is covered at the SDK wire boundary.
Filtering a `redacted_thinking`-only turn emits empty assistant content
instead of the converter's non-empty sentinel. Other request-policy
callers retain their previous behavior.

src/free_claude_code/providers/deepseek/compat.py

<details><summary><h3><a href="https://www.greptile.com/trex"><img
alt="T-Rex"
src="https://greptile-static-assets.s3.amazonaws.com/trex/trex_green.svg"
height="20" align="absmiddle"></a> T-Rex Logs</h3></summary>

**What T-Rex did**
- T-Rex executed the requested verification of the code base.
- T-Rex ran the focused test
wire\_messages\_keep\_prefix\_across\_tool\_thinking\_fallback with
pytest; the exact command and harness are captured in the log, and the
test completed with exit code 0.

<a
href="https://app.greptile.com/trex/runs/14533439/artifacts"><picture><source
media="(prefers-color-scheme: dark)"
srcset="https://greptile-static-assets.s3.amazonaws.com/badges/ViewAllArtifactsDark.svg?v=4"><source
media="(prefers-color-scheme: light)"
srcset="https://greptile-static-assets.s3.amazonaws.com/badges/ViewAllArtifacts.svg?v=4"><img
alt="View all artifacts"
src="https://greptile-static-assets.s3.amazonaws.com/badges/ViewAllArtifacts.svg?v=4"></picture></a>

<sub><a href="https://www.greptile.com/trex"><img alt="T-Rex"
src="https://greptile-static-assets.s3.amazonaws.com/trex/trex_green.svg"
height="14" align="absmiddle"></a> Ran code and verified through
T-Rex</sub>
</details>

<details open><summary><h3>Important Files Changed</h3></summary>

| Filename | Overview |
|----------|----------|
| src/free_claude_code/providers/deepseek/compat.py | Separates
historical tool reasoning from current thinking, but can serialize an
empty assistant message after filtering. |
| src/free_claude_code/providers/openai_chat/request_policy.py | Adds an
optional reasoning-history switch while preserving existing defaults for
other callers. |
| tests/providers/test_deepseek.py | Adds SDK wire-format tests for
stable prefixes and independent tool-reasoning replay. |
| pyproject.toml | Bumps the package patch version to 4.6.4. |
| uv.lock | Synchronizes the editable package version with the project
metadata. |
| ARCHITECTURE.md | Documents DeepSeek's per-turn reasoning replay and
append-only prefix behavior. |

</details>

<details open><summary><h3>Flowchart</h3></summary>

<a href="#gh-light-mode-only">

```mermaid
%%{init: {'theme': 'neutral'}}%%
flowchart LR
A[Anthropic history] --> B[Classify tool-call reasoning]
B --> C[Sanitize assistant history]
C --> D[Convert to OpenAI messages]
D --> E[Serialize through SDK]
E --> F[DeepSeek request]
B -->|Every tool call replayable| G[Keep current thinking]
B -->|Any tool call not replayable| H[Disable current thinking]
C -->|Tool-call turn| I[Retain reasoning]
C -->|Non-tool turn| J[Omit reasoning]
```

</a>
<a href="#gh-dark-mode-only">

```mermaid
%%{init: {'theme': 'base', 'themeVariables': {"darkMode": true, "background": "#0d1117", "primaryColor": "#21262d", "primaryTextColor": "#e6edf3", "primaryBorderColor": "#8b949e", "lineColor": "#8b949e", "textColor": "#e6edf3", "edgeLabelBackground": "#161b22", "actorBkg": "#21262d", "actorBorder": "#8b949e", "actorTextColor": "#e6edf3", "actorLineColor": "#8b949e", "signalColor": "#8b949e", "signalTextColor": "#e6edf3", "noteBkgColor": "#373320", "noteBorderColor": "#d4a72c", "noteTextColor": "#f0e6c0", "labelBoxBkgColor": "#21262d", "labelBoxBorderColor": "#8b949e", "labelTextColor": "#e6edf3", "loopTextColor": "#e6edf3", "activationBkgColor": "#30363d", "activationBorderColor": "#8b949e"}}}%%
flowchart LR
A[Anthropic history] --> B[Classify tool-call reasoning]
B --> C[Sanitize assistant history]
C --> D[Convert to OpenAI messages]
D --> E[Serialize through SDK]
E --> F[DeepSeek request]
B -->|Every tool call replayable| G[Keep current thinking]
B -->|Any tool call not replayable| H[Disable current thinking]
C -->|Tool-call turn| I[Retain reasoning]
C -->|Non-tool turn| J[Omit reasoning]
```

</a>
</details>

<a
href="https://app.greptile.com/api/ide/codex?prompt=IMPORTANT%3A%20Work%20in%20the%20repository%20%22alishahryar1%2Ffree-claude-code%22%20on%20the%20existing%20branch%20%22ali%2Fdeepseek-cache-stable-history%22.%20Checkout%20that%20branch%20%E2%80%94%20do%20NOT%20create%20a%20new%20branch%20or%20open%20a%20new%20PR.%20Push%20your%20changes%20to%20%22ali%2Fdeepseek-cache-stable-history%22.%0A%0AFix%20the%20following%201%20code%20review%20issue.%20Work%20through%20them%20one%20at%20a%20time%2C%20proposing%20concise%20fixes.%0A%0A---%0A%0A%23%23%23%20Issue%201%20of%201%0Asrc%2Ffree_claude_code%2Fproviders%2Fdeepseek%2Fcompat.py%3A146%0A**Empty%20Assistant%20Content%20Reaches%20Wire**%0A%0AWhen%20an%20assistant%20turn%20contains%20only%20%60redacted_thinking%60%2C%20filtering%20removes%20every%20block%20and%20replaces%20the%20list%20with%20%60%22%22%60.%20Revalidation%20then%20takes%20the%20string-content%20path%20and%20bypasses%20the%20converter's%20existing%20%60%22%20%22%60%20fallback%2C%20so%20DeepSeek%20can%20reject%20the%20continued%20conversation%20as%20an%20empty%20assistant%20message.%0A%0A%60%60%60suggestion%0A%20%20%20%20%20%20%20%20new_msg%5B%22content%22%5D%20%3D%20filtered%20or%20%22%20%22%0A%60%60%60%0A%0A&repo=alishahryar1%2Ffree-claude-code&pr=1126&platform=github"><picture><source
media="(prefers-color-scheme: dark)"
srcset="https://greptile-static-assets.s3.amazonaws.com/badges/FixAllInCodexDark.svg?v=6"><source
media="(prefers-color-scheme: light)"
srcset="https://greptile-static-assets.s3.amazonaws.com/badges/FixAllInCodex.svg?v=6"><img
alt="Fix All in Codex"
src="https://greptile-static-assets.s3.amazonaws.com/badges/FixAllInCodex.svg?v=6"></picture></a>

<sub>Reviews (1): Last reviewed commit: ["Preserve DeepSeek prompt cache
prefixes"](https://github.com/alishahryar1/free-claude-code/commit/b07ecf28a193e9cff1624e4ef75f38344a146d49)
| [Re-trigger
Greptile](https://app.greptile.com/api/retrigger?id=44440120)</sub>

> Greptile also left **1 inline comment** on this PR.

<!-- /greptile_comment -->
2026-07-15 09:38:09 -07:00
Ali Khokhar f0b31065ee Preserve mid-conversation system messages through provider conversion (#1125)
## Problem

FCC hoisted inline Anthropic `system` messages into the top-level system
prompt during request validation. Mid-conversation system messages are
position-sensitive, so this applied later instructions retroactively,
changed the existing prompt/cache prefix, and prevented provider
conversion from seeing the original transcript.

## Changes

- Preserve inline `system` messages, content, metadata, and ordering in
Messages and token-count requests while keeping the top-level system
prompt distinct.
- Convert text-only inline system messages to OpenAI Chat `system`
messages at the same transcript position; reject unrepresentable inline
blocks before streaming instead of silently dropping them.
- Remove the lossy normalization path and its unused role enum, and
document protocol-model versus target-conversion ownership in
`ARCHITECTURE.md`.
- Cover API routing, model serialization, cache-prefix stability, text
blocks, tool-result ordering, invalid content, and token counting; bump
the package to `4.6.2`.
- Verify all five local CI checks (2,287 tests) and the ordered
transcript against NVIDIA NIM, OpenRouter, Gemini, DeepSeek, Mistral,
and Hugging Face.

<!-- greptile_comment -->

<details open><summary><h3>Greptile Summary</h3></summary>

This PR preserves inline Anthropic system messages through provider
conversion. The main changes are:

- Keeps top-level and inline system content separate and ordered.
- Converts text-only inline system messages without moving them.
- Rejects system blocks that OpenAI Chat cannot represent safely.
- Updates request detection to ignore system context when counting user
turns.
- Adds serialization, routing, token-counting, and conversion coverage.
- Updates the package version and architecture documentation.
</details>

<h3>Confidence Score: 5/5</h3>

This looks safe to merge.

The leading-system detection path ignores system entries when counting
user turns. Inline system content remains ordered for provider
conversion. Unsupported system blocks fail explicitly instead of being
dropped. No blocking issues were found in the changed code.

No files require attention.

<details><summary><h3><a href="https://www.greptile.com/trex"><img
alt="T-Rex"
src="https://greptile-static-assets.s3.amazonaws.com/trex/trex_green.svg"
height="20" align="absmiddle"></a> T-Rex Logs</h3></summary>

**What T-Rex did**
- Validated that the transcript roles now follow the order user,
assistant, system, user and that the top-level prompt remains separate.
- Verified that inline system content is no longer counted in message
tokens and that cache\_control metadata survives parsing.
- Confirmed that the converted OpenAI transcript preserves position and
cache prefix.
- Observed that a system message following a tool result is converted as
assistant, tool, system.
- Ran the focused pytest and confirmed 209 passed in 2.82s with exit
code 0.

<a
href="https://app.greptile.com/trex/runs/14526513/artifacts"><picture><source
media="(prefers-color-scheme: dark)"
srcset="https://greptile-static-assets.s3.amazonaws.com/badges/ViewAllArtifactsDark.svg?v=4"><source
media="(prefers-color-scheme: light)"
srcset="https://greptile-static-assets.s3.amazonaws.com/badges/ViewAllArtifacts.svg?v=4"><img
alt="View all artifacts"
src="https://greptile-static-assets.s3.amazonaws.com/badges/ViewAllArtifacts.svg?v=4"></picture></a>

<sub><a href="https://www.greptile.com/trex"><img alt="T-Rex"
src="https://greptile-static-assets.s3.amazonaws.com/trex/trex_green.svg"
height="14" align="absmiddle"></a> Ran code and verified through
T-Rex</sub>
</details>

<details open><summary><h3>Important Files Changed</h3></summary>

| Filename | Overview |
|----------|----------|
| src/free_claude_code/core/anthropic/models.py | Preserves system-role
messages in the original transcript instead of hoisting them into the
top-level prompt. |
| src/free_claude_code/core/anthropic/conversion.py | Converts ordered
text-only system messages and rejects unsupported system content before
streaming. |
| src/free_claude_code/api/detection.py | Builds a read-only semantic
view of system context and conversational user turns for local request
detection. |

</details>

<sub>Reviews (2): Last reviewed commit: ["Restore optimizations with
inline
system..."](https://github.com/alishahryar1/free-claude-code/commit/6605ede7f604053381552106489dbd16bcd37987)
| [Re-trigger
Greptile](https://app.greptile.com/api/retrigger?id=44423885)</sub>

<!-- /greptile_comment -->
2026-07-15 03:37:11 -07:00
Ali Khokhar a092455b54 Add searchable model selection to the Admin UI (#1121)
## Problem

Admin model routing fields required users to construct provider-prefixed
model slugs. Optional tier overrides represented inheritance as an
unexplained blank value.

## Changes

| Before | After |
| --- | --- |
| Model inputs only gained suggestions after an individual provider
refresh. | Model inputs load configured and discovered canonical slugs
from one Admin catalog. |
| Model routing looked like unrestricted text entry. | Model routing
uses the browser's searchable model dropdown while retaining manual
entry. |
| Tier overrides displayed an empty value for fallback routing. | Tier
overrides display **None** and persist it as an unset override. |
| Model refresh returned provider-shaped cache internals. | Model
refresh returns the same canonical catalog consumed by the Admin UI. |
| Users inferred the provider/model slug format from examples. | The
Admin UI and README define and present complete provider/model slugs. |

<!-- greptile_comment -->

<details open><summary><h3>Greptile Summary</h3></summary>

This PR adds searchable model selection to the Admin UI. The main
changes are:

- Adds a canonical catalog of configured and discovered model slugs.
- Adds searchable model inputs while preserving manual entry.
- Represents unset tier overrides as **None**.
- Reports provider-specific model refresh failures.
- Reconciles cached models when provider settings change.
- Updates documentation, package metadata, and tests.
</details>

<h3>Confidence Score: 5/5</h3>

This looks safe to merge.

Catalog failures no longer stop the rest of the Admin UI from loading.
Partial provider refreshes now produce a visible warning. Removed
credential-backed providers are pruned from the shared cache, and later
stale writes are rejected. No blocking issues remain in the changed
code.

<details><summary><h3><a href="https://www.greptile.com/trex"><img
alt="T-Rex"
src="https://greptile-static-assets.s3.amazonaws.com/trex/trex_green.svg"
height="20" align="absmiddle"></a> T-Rex Logs</h3></summary>

**What T-Rex did**
- T-Rex ran the requested verification for the pull request checks.
- The verification completed, but local artifact references were not
uploaded.

<sub><a href="https://www.greptile.com/trex"><img alt="T-Rex"
src="https://greptile-static-assets.s3.amazonaws.com/trex/trex_green.svg"
height="14" align="absmiddle"></a> Ran code and verified through
T-Rex</sub>
</details>

<details open><summary><h3>Important Files Changed</h3></summary>

| Filename | Overview |
|----------|----------|
| src/free_claude_code/api/admin_static/admin.js | Adds searchable model
fields, optional catalog hydration, refresh warnings, and None-to-unset
conversion. |
| src/free_claude_code/api/admin_routes.py | Adds canonical model
catalog endpoints and provider refresh failure metadata. |
| src/free_claude_code/providers/runtime/discovery.py | Tracks provider
refresh outcomes and separates cache eligibility from discovery
eligibility. |
| src/free_claude_code/providers/runtime/model_cache.py | Scopes cached
model metadata to currently available providers and removes stale remote
entries. |
| src/free_claude_code/runtime/provider_manager.py | Reconciles cache
scope during runtime replacement and returns explicit refresh results. |

</details>

<sub>Reviews (2): Last reviewed commit: ["Fix model catalog refresh
lifecycle"](https://github.com/alishahryar1/free-claude-code/commit/d16e170055f5389e538e18dece269a5f7a8c599d)
| [Re-trigger
Greptile](https://app.greptile.com/api/retrigger?id=44373795)</sub>

<!-- /greptile_comment -->
2026-07-15 01:27:36 -07:00
Ali Khokhar b1877a4b21 Add Ollama Cloud as a first-class provider (#1106)
## Problem

FCC supports Ollama only through a local daemon, so users cannot
authenticate directly to Ollama Cloud or discover its hosted models from
the Admin UI. Ollama's OpenAI-compatible API also uses the standard
`reasoning` field instead of `reasoning_content`, which would otherwise
drop thinking output and tool-history reasoning.

## Changes

| Before | After |
| --- | --- |
| `ollama/...` requires a local Ollama server. | `ollama_cloud/...`
connects directly to `https://ollama.com/v1` with `OLLAMA_API_KEY`,
while local Ollama remains unchanged. |
| OpenAI-chat reasoning was hard-coded to `reasoning_content`. |
Provider profiles declare their reasoning field, so Ollama streams and
replays `reasoning` without a specialized transport. |
| Ollama Cloud was absent from configuration, model discovery, docs, and
smoke coverage. | The catalog, Admin UI, proxy setting, model picker,
README, smoke matrix, and `4.5.0` release metadata expose the provider
consistently. |

<!-- greptile_comment -->

<details open><summary><h3>Greptile Summary</h3></summary>

This PR adds Ollama Cloud as a separate OpenAI-compatible provider. The
main changes are:

- New `ollama_cloud` catalog entry, settings, admin fields, proxy
setting, and smoke configuration.
- Provider profiles now choose the streamed and replayed reasoning field
per provider.
- Ollama Cloud uses `reasoning` and `reasoning_effort`, while local
Ollama stays on its separate local configuration.
- Streaming recovery now respects the resolved thinking setting when
collecting reasoning.
- Tests, docs, environment examples, and release metadata were updated
for the new provider.
</details>

<h3>Confidence Score: 5/5</h3>

This looks safe to merge.

No blocking issues found in the changed code.

No files need attention.

<details><summary><h3><a href="https://www.greptile.com/trex"><img
alt="T-Rex"
src="https://greptile-static-assets.s3.amazonaws.com/trex/trex_green.svg"
height="20" align="absmiddle"></a> T-Rex Logs</h3></summary>

**What T-Rex did**
- Ran the provider/runtime/converter/streaming tests with full verbose
pytest output and confirmed EXIT\_CODE: 0.
- Ran the config/catalog/contracts/admin tests with full verbose pytest
output and confirmed EXIT\_CODE: 0.
- Generated and ran an introspection script to verify catalog, profile,
settings, and admin manifest values without external API calls, and
recorded the local execution output.

<a
href="https://app.greptile.com/trex/runs/14367965/artifacts"><picture><source
media="(prefers-color-scheme: dark)"
srcset="https://greptile-static-assets.s3.amazonaws.com/badges/ViewAllArtifactsDark.svg?v=4"><source
media="(prefers-color-scheme: light)"
srcset="https://greptile-static-assets.s3.amazonaws.com/badges/ViewAllArtifacts.svg?v=4"><img
alt="View all artifacts"
src="https://greptile-static-assets.s3.amazonaws.com/badges/ViewAllArtifacts.svg?v=4"></picture></a>

<sub><a href="https://www.greptile.com/trex"><img alt="T-Rex"
src="https://greptile-static-assets.s3.amazonaws.com/trex/trex_green.svg"
height="14" align="absmiddle"></a> Ran code and verified through
T-Rex</sub>
</details>

<details open><summary><h3>Important Files Changed</h3></summary>

| Filename | Overview |
|----------|----------|
| src/free_claude_code/providers/openai_chat/profiles.py | Adds
provider-level reasoning field selection and the Ollama Cloud profile. |
| src/free_claude_code/providers/openai_chat/provider.py | Uses the
profile reasoning field for streaming and passes the thinking setting
into recovery. |
| src/free_claude_code/providers/openai_chat/request_policy.py | Selects
reasoning replay mode from the provider policy only when thinking is
enabled. |
| src/free_claude_code/core/anthropic/conversion.py | Supports replaying
assistant reasoning through either `reasoning_content` or `reasoning`. |
| src/free_claude_code/config/provider_catalog.py | Registers Ollama
Cloud as a remote provider distinct from local Ollama. |
| src/free_claude_code/config/settings.py | Adds the Ollama Cloud API
key and proxy settings. |

</details>

<sub>Reviews (3): Last reviewed commit: ["Keep local Ollama wire
behavior
unchange..."](https://github.com/alishahryar1/free-claude-code/commit/d6f97cdb0074a4cbda140d484fb6dd679a1406df)
| [Re-trigger
Greptile](https://app.greptile.com/api/retrigger?id=44089262)</sub>

<!-- /greptile_comment -->
2026-07-14 02:48:03 -07:00
Ali Khokhar 4f56a6aa17 Add Fable as a first-class Claude routing tier (#1099)
## Problem

Claude Code now sends `claude-fable-5` for the Fable alias, but FCC
treated it as an unrecognized model and collapsed it into the global
fallback route. Users could not map Fable traffic or reasoning behavior
independently. Fixes #1097.

## Changes

| Before | After |
| --- | --- |
| Fable requests inherited `MODEL` and `ENABLE_MODEL_THINKING`. | Fable
requests use `MODEL_FABLE` and `ENABLE_FABLE_THINKING` when configured,
otherwise inherit the existing defaults. |
| `/v1/models` omitted Claude Fable 5. | `/v1/models` advertises the
canonical `claude-fable-5` identifier. |
| Admin, documentation, validation, and smoke contracts described three
Claude tiers. | Admin, documentation, validation, and smoke contracts
describe Fable alongside Opus, Sonnet, and Haiku. |
| The package version was `4.3.1`. | The package version is `4.4.0`. |

<!-- greptile_comment -->

<details open><summary><h3>Greptile Summary</h3></summary>

This PR adds Fable as a Claude routing tier. The main changes are:

- `MODEL_FABLE` and `ENABLE_FABLE_THINKING` settings.
- Fable routing and thinking resolution in `ModelRouter`.
- `claude-fable-5` in the model catalog.
- Admin, docs, smoke, and test coverage updates.
- Package version bump to `4.4.0`.
</details>

<h3>Confidence Score: 4/5</h3>

The changed routing path needs a fix for direct provider model ids
containing `fable`.

Fable settings, validation, admin fields, and model listing are
consistent with the existing tier patterns. Blank Fable settings inherit
the existing defaults. Direct provider model ids can receive the Fable
thinking override when their model name contains `fable`.

src/free_claude_code/application/routing.py

<details><summary><h3><a href="https://www.greptile.com/trex"><img
alt="T-Rex"
src="https://greptile-static-assets.s3.amazonaws.com/trex/trex_green.svg"
height="20" align="absmiddle"></a> T-Rex Logs</h3></summary>

**What T-Rex did**
- Reproduced the Fable thinking overmatch by running a focused Python
repro that disables global thinking and enables Fable thinking, then
resolves sambanova/my-fable-ensemble-v2.
- The repro confirmed direct routing preserved provider\_id=sambanova,
provider\_model=my-fable-ensemble-v2, and
provider\_model\_ref=sambanova/my-fable-ensemble-v2, with
resolved\_thinking\_enabled and thinking\_enabled both true.
- Ran the fable-tier validation pytest, which finished with exit code 0
and 177 tests passed.
- Ran the runtime probe to exercise the API model list, Settings env
parsing, and ModelRouter paths, and observed a 200 OK on GET /v1/models,
with the catalog item id claude-fable-5 and the expected Fable vs
global-default routing behavior.
- Generated the probe source file used to exercise the API and Settings
paths, enabling repeatable validation without real provider credentials.

<a
href="https://app.greptile.com/trex/runs/14275379/artifacts"><picture><source
media="(prefers-color-scheme: dark)"
srcset="https://greptile-static-assets.s3.amazonaws.com/badges/ViewAllArtifactsDark.svg?v=4"><source
media="(prefers-color-scheme: light)"
srcset="https://greptile-static-assets.s3.amazonaws.com/badges/ViewAllArtifacts.svg?v=4"><img
alt="View all artifacts"
src="https://greptile-static-assets.s3.amazonaws.com/badges/ViewAllArtifacts.svg?v=4"></picture></a>

<sub><a href="https://www.greptile.com/trex"><img alt="T-Rex"
src="https://greptile-static-assets.s3.amazonaws.com/trex/trex_green.svg"
height="14" align="absmiddle"></a> Ran code and verified through
T-Rex</sub>
</details>

<details open><summary><h3>Important Files Changed</h3></summary>

| Filename | Overview |
|----------|----------|
| src/free_claude_code/application/routing.py | Adds Fable model and
thinking branches; the thinking branch can also match unrelated direct
provider model ids containing `fable`. |
| src/free_claude_code/config/settings.py | Adds optional Fable model
and thinking settings with blank-env inheritance and provider/model
validation. |
| src/free_claude_code/config/model_refs.py | Includes Fable in
configured chat model reference collection and dedupe. |
| src/free_claude_code/config/admin/manifest.py | Adds Fable model and
thinking controls to the admin manifest. |
| src/free_claude_code/api/model_catalog.py | Adds `claude-fable-5` to
the advertised Claude model aliases. |
| pyproject.toml | Bumps the package version to `4.4.0`. |
| uv.lock | Updates the editable package version to match
`pyproject.toml`. |

</details>

<details open><summary><h3>Flowchart</h3></summary>

<a href="#gh-light-mode-only">

```mermaid
%%{init: {'theme': 'neutral'}}%%
flowchart TD
  A[Incoming model name] --> B{Direct provider or gateway id?}
  B -- yes --> C[Use provider/model directly]
  C --> D[Resolve thinking from provider model string]
  B -- no --> E{Claude tier match}
  E -- Fable --> F[MODEL_FABLE or MODEL]
  E -- Opus/Sonnet/Haiku --> G[Tier override or MODEL]
  E -- None --> H[MODEL]
  D --> I[Provider request]
  F --> I
  G --> I
  H --> I
```

</a>
<a href="#gh-dark-mode-only">

```mermaid
%%{init: {'theme': 'base', 'themeVariables': {"darkMode": true, "background": "#0d1117", "primaryColor": "#21262d", "primaryTextColor": "#e6edf3", "primaryBorderColor": "#8b949e", "lineColor": "#8b949e", "textColor": "#e6edf3", "edgeLabelBackground": "#161b22", "actorBkg": "#21262d", "actorBorder": "#8b949e", "actorTextColor": "#e6edf3", "actorLineColor": "#8b949e", "signalColor": "#8b949e", "signalTextColor": "#e6edf3", "noteBkgColor": "#373320", "noteBorderColor": "#d4a72c", "noteTextColor": "#f0e6c0", "labelBoxBkgColor": "#21262d", "labelBoxBorderColor": "#8b949e", "labelTextColor": "#e6edf3", "loopTextColor": "#e6edf3", "activationBkgColor": "#30363d", "activationBorderColor": "#8b949e"}}}%%
flowchart TD
  A[Incoming model name] --> B{Direct provider or gateway id?}
  B -- yes --> C[Use provider/model directly]
  C --> D[Resolve thinking from provider model string]
  B -- no --> E{Claude tier match}
  E -- Fable --> F[MODEL_FABLE or MODEL]
  E -- Opus/Sonnet/Haiku --> G[Tier override or MODEL]
  E -- None --> H[MODEL]
  D --> I[Provider request]
  F --> I
  G --> I
  H --> I
```

</a>
</details>

<a
href="https://app.greptile.com/api/ide/codex?prompt=IMPORTANT%3A%20Work%20in%20the%20repository%20%22alishahryar1%2Ffree-claude-code%22%20on%20the%20existing%20branch%20%22ali%2Fadd-fable-routing-tier%22.%20Checkout%20that%20branch%20%E2%80%94%20do%20NOT%20create%20a%20new%20branch%20or%20open%20a%20new%20PR.%20Push%20your%20changes%20to%20%22ali%2Fadd-fable-routing-tier%22.%0A%0AFix%20the%20following%201%20code%20review%20issue.%20Work%20through%20them%20one%20at%20a%20time%2C%20proposing%20concise%20fixes.%0A%0A---%0A%0A%23%23%23%20Issue%201%20of%201%0Asrc%2Ffree_claude_code%2Fapplication%2Frouting.py%3A134-135%0A**Fable%20Thinking%20Overmatches%20Models**%0A%0AWhen%20a%20direct%20provider%20request%20uses%20a%20model%20id%20like%20%60sambanova%2Fmy-fable-ensemble-v2%60%2C%20the%20direct%20route%20bypasses%20tier%20remapping%20but%20still%20calls%20%60_resolve_thinking%28%29%60%20with%20the%20provider%20model%20string.%20With%20%60ENABLE_FABLE_THINKING%60%20set%2C%20this%20substring%20check%20applies%20Fable%20thinking%20behavior%20to%20an%20unrelated%20provider%20model%2C%20changing%20the%20outgoing%20request%20shape%20just%20because%20the%20model%20id%20contains%20%60fable%60.%0A%0A&repo=alishahryar1%2Ffree-claude-code&pr=1099&platform=github"><picture><source
media="(prefers-color-scheme: dark)"
srcset="https://greptile-static-assets.s3.amazonaws.com/badges/FixAllInCodexDark.svg?v=6"><source
media="(prefers-color-scheme: light)"
srcset="https://greptile-static-assets.s3.amazonaws.com/badges/FixAllInCodex.svg?v=6"><img
alt="Fix All in Codex"
src="https://greptile-static-assets.s3.amazonaws.com/badges/FixAllInCodex.svg?v=6"></picture></a>

<sub>Reviews (1): Last reviewed commit: ["Add Fable as a first-class
routing
tier"](https://github.com/alishahryar1/free-claude-code/commit/0705840c65511fd85c74fd2c62ba8ea97afe7c12)
| [Re-trigger
Greptile](https://app.greptile.com/api/retrigger?id=43892323)</sub>

> Greptile also left **1 inline comment** on this PR.

<!-- /greptile_comment -->
2026-07-13 11:56:34 -07:00
Fethi Amari 1b4c7e6779 feat: add vision support for OpenAI chat conversion (#1061)
## Problem

FCC rejects Anthropic user image blocks before vision-capable
OpenAI-compatible providers can receive them. This blocks pasted images
in Claude Code and IDE clients. Fixes #512.

## Changes

| Before | After |
| --- | --- |
| User image blocks fail during Anthropic-to-OpenAI conversion. | Base64
and URL image sources become OpenAI `image_url` parts. |
| Mixed text and image input cannot reach an upstream model. | Mixed
content preserves exact block order and text-only behavior. |
| Invalid image sources require ambiguous downstream handling. | Missing
or unsupported source fields fail explicitly before execution. |
| Vision support could require flags or proxy-side downloads. |
Conversion stays provider-neutral and performs no network I/O. |
| The package remains on `4.1.0`. | The package advances to `4.2.0` with
a regenerated lockfile. |
| Converter coverage rejects every user image. | Tests cover source
mapping, ordering, tool boundaries, validation, and request
construction. |

<!-- greptile_comment -->

<h3>Greptile Summary</h3>

This PR adds OpenAI-compatible vision conversion for Anthropic image
blocks. The main changes are:

- Converts base64 image sources into data image URLs.
- Converts URL image sources into OpenAI `image_url` parts.
- Preserves mixed text and image ordering in user messages.
- Rejects unsupported or incomplete image sources.
- Updates converter tests for image and tool-result cases.
- Bumps package metadata to `4.2.0` in both project and lock files.

<h3>Confidence Score: 5/5</h3>

This looks safe to merge.

No blocking issues found in the changed code.

No files need attention.

<details><summary><h3><a href="https://www.greptile.com/trex"><img
alt="T-Rex"
src="https://greptile-static-assets.s3.amazonaws.com/trex/trex_green.svg"
height="20" align="absmiddle"></a> T-Rex Logs</h3></summary>

**What T-Rex did**
- Reviewed the initial focused run and confirmed all tests passed, with
the wrapper exit code reported as nonzero and retained for traceability.
- Validated the focused rerun completed cleanly with 10 tests passed and
exit code 0.
- Validated the full converter run completed with 73 tests passed and
exit code 0.

<a
href="https://app.greptile.com/trex/runs/14209442/artifacts"><picture><source
media="(prefers-color-scheme: dark)"
srcset="https://greptile-static-assets.s3.amazonaws.com/badges/ViewAllArtifactsDark.svg?v=4"><source
media="(prefers-color-scheme: light)"
srcset="https://greptile-static-assets.s3.amazonaws.com/badges/ViewAllArtifacts.svg?v=4"><img
alt="View all artifacts"
src="https://greptile-static-assets.s3.amazonaws.com/badges/ViewAllArtifacts.svg?v=4"></picture></a>

<sub><a href="https://www.greptile.com/trex"><img alt="T-Rex"
src="https://greptile-static-assets.s3.amazonaws.com/trex/trex_green.svg"
height="14" align="absmiddle"></a> Ran code and verified through
T-Rex</sub>
</details>

<h3>Important Files Changed</h3>

| Filename | Overview |
|----------|----------|
| src/free_claude_code/core/anthropic/conversion.py | Adds ordered image
conversion for OpenAI-compatible chat content. |
| tests/providers/test_converter.py | Adds tests for image conversion,
ordering, validation, and request-body output. |
| pyproject.toml | Updates the package version to `4.2.0`. |
| uv.lock | Updates the locked local package version to `4.2.0`. |
| ARCHITECTURE.md | Documents core image conversion behavior for
OpenAI-compatible providers. |
| README.md | Mentions image input support across compatible models. |

<sub>Reviews (8): Last reviewed commit: ["feat: add vision support for
OpenAI
chat..."](https://github.com/alishahryar1/free-claude-code/commit/4a6b6de14d19ce22e3be5981f84bb6f236a11181)
| [Re-trigger
Greptile](https://app.greptile.com/api/retrigger?id=43544835)</sub>

<!-- /greptile_comment -->
2026-07-13 03:16:56 -07:00
Ali Khokhar f2f5b08714 Make provider retry tests deterministic (#1070)
## Problem

Provider retry integration tests mocked `asyncio.sleep` while reactive
deadlines still used `time.monotonic`. Serial runs busy-waited through
real backoffs and could exceed validation timeouts despite passing.

## Changes

| Before | After |
| --- | --- |
| Tests globally mocked sleep without advancing the limiter clock. | A
deterministic test limiter runs real retry classification and attempt
accounting with zero delay. |
| Provider integration tests also exercised reactive timing already
covered by limiter tests. | Provider integration tests focus on retry
and terminal-error contracts; limiter timing remains in its dedicated
suite. |
| Serial retry validation could take minutes. | The complete retry file
finishes in about two seconds including startup. |

<!-- greptile_comment -->

<details open><summary><h3>Greptile Summary</h3></summary>

This PR makes provider retry tests run without wall-clock backoff. The
main changes are:

- Adds an immediate retry limiter for provider tests.
- Forces retry backoff parameters to zero in that test limiter.
- Skips reactive block timing in provider retry tests.
- Removes broad `asyncio.sleep` mocks from the OpenAI-compatible retry
tests.
</details>

<h3>Confidence Score: 5/5</h3>

This looks safe to merge.

No blocking issues found in the changed code.

None.

<details><summary><h3><a href="https://www.greptile.com/trex"><img
alt="T-Rex"
src="https://greptile-static-assets.s3.amazonaws.com/trex/trex_green.svg"
height="20" align="absmiddle"></a> T-Rex Logs</h3></summary>

**What T-Rex did**
- T-Rex executed the provider retry test and captured a log showing the
exact command, working directory, timestamp, Pytest output, elapsed
time, and exit code.

<a
href="https://app.greptile.com/trex/runs/14126948/artifacts"><picture><source
media="(prefers-color-scheme: dark)"
srcset="https://greptile-static-assets.s3.amazonaws.com/badges/ViewAllArtifactsDark.svg?v=4"><source
media="(prefers-color-scheme: light)"
srcset="https://greptile-static-assets.s3.amazonaws.com/badges/ViewAllArtifacts.svg?v=4"><img
alt="View all artifacts"
src="https://greptile-static-assets.s3.amazonaws.com/badges/ViewAllArtifacts.svg?v=4"></picture></a>

<sub><a href="https://www.greptile.com/trex"><img alt="T-Rex"
src="https://greptile-static-assets.s3.amazonaws.com/trex/trex_green.svg"
height="14" align="absmiddle"></a> Ran code and verified through
T-Rex</sub>
</details>

<details open><summary><h3>Important Files Changed</h3></summary>

| Filename | Overview |
|----------|----------|
| tests/providers/support.py | Adds a test-only limiter that keeps retry
classification and attempt accounting while removing elapsed-time
behavior. |
| tests/providers/test_openai_compat_5xx_retry.py | Removes sleep mocks
from retry tests that now rely on the immediate retry limiter. |

</details>

<sub>Reviews (1): Last reviewed commit: ["Make provider retry tests
deterministic"](https://github.com/alishahryar1/free-claude-code/commit/e5f8c05ec5900586df8d4e1e7689288f38eeab15)
| [Re-trigger
Greptile](https://app.greptile.com/api/retrigger?id=43584603)</sub>

<!-- /greptile_comment -->
2026-07-11 20:23:40 -07:00
Ali Khokhar f3ea35a777 Consolidate OpenAI-compatible provider adapters (#1069)
## Problem

Sixteen OpenAI-compatible providers were represented by
configuration-only subclasses and one factory function each. The sole
provider family also lived under a multi-transport namespace that no
longer described the codebase, while provider defaults, IDs, and
instance caching passed through redundant forwarding layers.

## Changes

| Before | After |
| --- | --- |
| Sixteen provider IDs used configuration-only subclasses. | Immutable
OpenAI-chat profiles configure one concrete provider while preserving
each provider's request policy. |
| Every provider required a dedicated factory function. | Generic
profile construction is the default; only eight adapters with real state
or algorithms retain factories. |
| The sole provider family lived under `providers/transports/`. |
`providers/openai_chat/` directly owns shared request, stream, recovery,
tool, and usage behavior. |
| Provider defaults and IDs passed through forwarding modules, and
constructors repeated fallback resolution. | The neutral catalog
resolves complete immutable provider configuration once. |
| `ProviderRuntime` wrapped a pass-through provider cache. |
`ProviderRuntime` directly owns lazy provider instances and cleanup. |
| Production carried the duplicated adapter structure. | The final shape
removes a net 873 production lines with no compatibility shim or
customer-facing provider change. |

<!-- greptile_comment -->

<details open><summary><h3>Greptile Summary</h3></summary>

This PR consolidates OpenAI-compatible providers behind shared profiles
and runtime construction. The main changes are:

- OpenAI-chat behavior moved into `providers/openai_chat`.
- Configuration-only provider subclasses replaced by immutable profiles.
- Provider runtime now owns lazy instance caching and cleanup directly.
- Provider defaults and IDs now resolve through the neutral catalog.
- Tests and smoke helpers updated for the new provider shape.
</details>


<h3>Confidence Score: 5/5</h3>

This looks safe to merge.

No blocking issues found in the changed code.

No files need attention.

<details><summary><h3><a href="https://www.greptile.com/trex"><img
alt="T-Rex"
src="https://greptile-static-assets.s3.amazonaws.com/trex/trex_green.svg"
height="20" align="absmiddle"></a> T-Rex Logs</h3></summary>

**What T-Rex did**
- T-Rex produced a proof for a posted P1 finding and linked it to the
corresponding review comment for details.
- T-Rex saved contract-validation logs for provider consolidation and
preserved four log files that show the test hanging boundary and the
timeout exit code 124.

<a
href="https://app.greptile.com/trex/runs/14125428/artifacts"><picture><source
media="(prefers-color-scheme: dark)"
srcset="https://greptile-static-assets.s3.amazonaws.com/badges/ViewAllArtifactsDark.svg?v=4"><source
media="(prefers-color-scheme: light)"
srcset="https://greptile-static-assets.s3.amazonaws.com/badges/ViewAllArtifacts.svg?v=4"><img
alt="View all artifacts"
src="https://greptile-static-assets.s3.amazonaws.com/badges/ViewAllArtifacts.svg?v=4"></picture></a>

<sub><a href="https://www.greptile.com/trex"><img alt="T-Rex"
src="https://greptile-static-assets.s3.amazonaws.com/trex/trex_green.svg"
height="14" align="absmiddle"></a> Ran code and verified through
T-Rex</sub>
</details>


<details open><summary><h3>Important Files Changed</h3></summary>




| Filename | Overview |
|----------|----------|
| src/free_claude_code/providers/runtime/config.py | Builds resolved
provider configuration from catalog descriptors, including static
credentials for local providers. |
| src/free_claude_code/providers/openai_chat/provider.py | Creates
shared OpenAI-chat clients from resolved configuration and immutable
provider profiles. |
| src/free_claude_code/providers/openai_chat/profiles.py | Defines
declarative profiles for formerly configuration-only OpenAI-compatible
adapters. |

</details>


<!-- greptile_failed_comments -->
<h3>Comments Outside Diff (2)</h3>

1. General comment 

<a href="#"><img alt="P1"
src="https://greptile-static-assets.s3.amazonaws.com/badges/p1.svg?v=9"
align="top"></a> **NIM OpenAI-compatible 504 exhausted streaming retry
path hangs instead of completing with the expected user-facing error**

   - **Bug**
- The broader provider consolidation suite failed all parametrized
`test_nim_stream_openai_5xx_exhausted_emits_user_message` cases under
xdist with worker crashes. A serial isolation run narrowed this to the
same OpenAI-compatible NIM exhausted 5xx streaming path: 500, 502, and
503 completed, but the 504 case did not finish before the explicit 120s
timeout, producing exit code 124. This contradicts the expected
streaming retry contract that exhausted transient provider failures
terminate and emit the configured user-facing error.
   - **Cause**
- The OpenAI-compatible streaming retry/exhaustion path for NIM 504
responses appears to wait indefinitely or otherwise fail to terminate
after retries are exhausted. The exact code-level loop/await point was
not isolated within the validation budget, but the failure is anchored
to the provider streaming retry contract exercised by
`tests/providers/test_openai_compat_5xx_retry.py::test_nim_stream_openai_5xx_exhausted_emits_user_message[504-temporarily
unavailable]`.
   - **Fix**
- Inspect the OpenAI-compatible/NIM streaming retry exhaustion handling
for 504 responses and ensure retry limits are enforced, the async stream
is closed/cancelled on exhaustion, and the provider raises/emits the
same terminal user-facing error contract as the 500/502/503 cases. Add
or keep a serial regression test for the 504 exhausted stream path to
prevent hangs.

<sub><a href="https://www.greptile.com/trex"><img alt="T-Rex"
src="https://greptile-static-assets.s3.amazonaws.com/trex/trex_green.svg"
height="14" align="absmiddle"></a> Ran code and verified through
T-Rex</sub>

2. General comment 

<a href="#"><img alt="P1"
src="https://greptile-static-assets.s3.amazonaws.com/badges/p1.svg?v=9"
align="top"></a> **OpenAI-compatible 5xx retry suite hangs on exhausted
502 retry path**

   - **Bug**
- The recommended non-xdist retry validation does not complete. A
verbose rerun with a 60 second timeout shows the suite passes tests
through the exhausted 500 case, then times out while running
`test_nim_stream_openai_5xx_exhausted_emits_user_message[502-temporarily
unavailable]`. This indicates the consolidated OpenAI-compatible
provider retry/error path can hang for at least the 502 exhausted-error
scenario, preventing reliable validation and potentially blocking
callers from receiving the expected `ExecutionFailure`.
   - **Cause**
- The OpenAI-compatible provider's exhausted 502 retry/error handling
path appears not to terminate promptly under the mocked repeated
`openai.InternalServerError` scenario. The exact code location was not
changed during validation, but the failure is isolated to the shared
OpenAI-chat 5xx retry behavior exercised by
`NvidiaNimProvider.stream_response`.
   - **Fix**
- Debug the shared OpenAI-chat retry loop/error classification for 502
responses. Ensure retries are bounded, patched `asyncio.sleep` is
awaited without real backoff during tests, and exhausted 502/503/504
errors consistently raise `ExecutionFailure` with the expected
temporary-unavailable message instead of continuing work indefinitely.

<sub><a href="https://www.greptile.com/trex"><img alt="T-Rex"
src="https://greptile-static-assets.s3.amazonaws.com/trex/trex_green.svg"
height="14" align="absmiddle"></a> Ran code and verified through
T-Rex</sub>

<!-- /greptile_failed_comments -->

<sub>Reviews (2): Last reviewed commit: ["Prove local provider
credential
resoluti..."](https://github.com/alishahryar1/free-claude-code/commit/ba226d97efdae1889bea0372c7bacf268122d65c)
| [Re-trigger
Greptile](https://app.greptile.com/api/retrigger?id=43580983)</sub>

<!-- /greptile_comment -->
2026-07-11 19:55:19 -07:00
Ali Khokhar 3a7e0ccf7a Remove the native Anthropic provider transport (#1067)
## Problem

Ollama and llama.cpp still used a parallel native Anthropic transport
after the other providers moved to OpenAI Chat. That kept duplicate
request, SSE, recovery, model-list, and server-tool policy machinery
alive.

## Changes

| Before | After |
| --- | --- |
| Ollama and llama.cpp streamed through provider-specific Anthropic
`/messages` adapters. | Ollama and llama.cpp use the shared OpenAI Chat
transport. |
| Native request serialization, SSE normalization, error mapping, and
recovery remained beside the OpenAI path. | Native-only machinery is
removed and all providers share one transport lifecycle. |
| Routed models carried a capability object solely to permit native
server-tool passthrough. | Routing carries only route decisions; FCC
handles forced server tools locally and rejects lossy passthrough. |
| Ollama discovery used a separate `/api/tags` parser and rejected `/v1`
configuration. | Ollama discovery uses `/v1/models` and accepts either
root or `/v1` base URLs. |
| Obsolete native tests and a compatibility facade kept deleted
internals represented. | Tests cover the shared transport and real
Ollama product path without compatibility shims. |

<!-- greptile_comment -->

<details open><summary><h3>Greptile Summary</h3></summary>

This PR removes the native Anthropic transport path for local providers.
The main changes are:

- Ollama and llama.cpp now use the shared OpenAI Chat transport.
- Local provider base URLs are normalized to the OpenAI-compatible `/v1`
API root.
- Ollama discovery now uses the OpenAI-compatible model-listing path.
- Native Anthropic transport code and server-tool passthrough capability
metadata were removed.
- Tests and smoke coverage were updated for the shared transport path.
</details>


<h3>Confidence Score: 5/5</h3>

This looks safe to merge.

No blocking issues found in the changed code.

None.

<details><summary><h3><a href="https://www.greptile.com/trex"><img
alt="T-Rex"
src="https://greptile-static-assets.s3.amazonaws.com/trex/trex_green.svg"
height="20" align="absmiddle"></a> T-Rex Logs</h3></summary>

**What T-Rex did**
- Ran the targeted local provider pytest slice and observed exit code 0.
- Executed the generated runtime harness to emulate the provider HTTP
interactions and capture a request trace.
- Validated the request trace showed two GET /v1/models calls authorized
as Bearer ollama for root and /v1 base URL configurations, and a POST
/v1/chat/completions authorized as Bearer llamacpp with streaming OpenAI
chat JSON payload.
- Confirmed the exact generated harness script used for the runtime
proof is the harness file referenced in the artifacts.

<a
href="https://app.greptile.com/trex/runs/14121848/artifacts"><picture><source
media="(prefers-color-scheme: dark)"
srcset="https://greptile-static-assets.s3.amazonaws.com/badges/ViewAllArtifactsDark.svg?v=4"><source
media="(prefers-color-scheme: light)"
srcset="https://greptile-static-assets.s3.amazonaws.com/badges/ViewAllArtifacts.svg?v=4"><img
alt="View all artifacts"
src="https://greptile-static-assets.s3.amazonaws.com/badges/ViewAllArtifacts.svg?v=4"></picture></a>

<sub><a href="https://www.greptile.com/trex"><img alt="T-Rex"
src="https://greptile-static-assets.s3.amazonaws.com/trex/trex_green.svg"
height="14" align="absmiddle"></a> Ran code and verified through
T-Rex</sub>
</details>


<details open><summary><h3>Important Files Changed</h3></summary>




| Filename | Overview |
|----------|----------|
| src/free_claude_code/providers/transports/openai_chat/base_url.py |
Adds a helper that normalizes local OpenAI-compatible server roots to
`/v1`. |
| src/free_claude_code/providers/llamacpp/client.py | Moves llama.cpp to
the shared OpenAI Chat transport with local base URL normalization. |
| src/free_claude_code/providers/ollama/client.py | Moves Ollama to the
shared OpenAI Chat transport with local base URL normalization. |
| src/free_claude_code/api/handlers/messages.py | Applies server-tool
rejection through the shared request policy instead of provider
passthrough metadata. |
| src/free_claude_code/application/routing.py | Removes provider
capability metadata from routed model results. |

</details>


<!-- greptile_failed_comments -->
<h3>Comments Outside Diff (1)</h3>

1. `src/free_claude_code/api/handlers/messages.py`, line 261-267
([link](https://github.com/alishahryar1/free-claude-code/blob/3c7ff176da46560c4d27b3846dca1ab1c7db561c/src/free_claude_code/api/handlers/messages.py#L261-L267))

<a href="#"><img alt="P2"
src="https://greptile-static-assets.s3.amazonaws.com/badges/p2.svg?v=9"
align="top"></a> **Native Server Tools Always Reject**

With the passthrough capability check removed, Ollama and llama.cpp
requests that previously used their native Anthropic transport for
`web_search` or `web_fetch` are rejected before provider execution. The
default `ENABLE_WEB_SERVER_TOOLS=false` now makes forced server-tool
requests return an invalid-request error instead of reaching the local
provider path that used to support them.

<a
href="https://app.greptile.com/api/ide/codex?prompt=IMPORTANT%3A%20Work%20in%20the%20repository%20%22alishahryar1%2Ffree-claude-code%22%20on%20the%20existing%20branch%20%22ali%2Fremove-native-anthropic-transport%22.%20Checkout%20that%20branch%20%E2%80%94%20do%20NOT%20create%20a%20new%20branch%20or%20open%20a%20new%20PR.%20Push%20your%20changes%20to%20%22ali%2Fremove-native-anthropic-transport%22.%0A%0AThis%20is%20a%20comment%20left%20during%20a%20code%20review.%0APath%3A%20src%2Ffree_claude_code%2Fapi%2Fhandlers%2Fmessages.py%0ALine%3A%20261-267%0A%0AComment%3A%0A**Native%20Server%20Tools%20Always%20Reject**%0A%0AWith%20the%20passthrough%20capability%20check%20removed%2C%20Ollama%20and%20llama.cpp%20requests%20that%20previously%20used%20their%20native%20Anthropic%20transport%20for%20%60web_search%60%20or%20%60web_fetch%60%20are%20rejected%20before%20provider%20execution.%20The%20default%20%60ENABLE_WEB_SERVER_TOOLS%3Dfalse%60%20now%20makes%20forced%20server-tool%20requests%20return%20an%20invalid-request%20error%20instead%20of%20reaching%20the%20local%20provider%20path%20that%20used%20to%20support%20them.%0A%0AHow%20can%20I%20resolve%20this%3F%20If%20you%20propose%20a%20fix%2C%20please%20make%20it%20concise.&repo=alishahryar1%2Ffree-claude-code&pr=1067&platform=github"><picture><source
media="(prefers-color-scheme: dark)"
srcset="https://greptile-static-assets.s3.amazonaws.com/badges/FixInCodexDark.svg?v=6"><source
media="(prefers-color-scheme: light)"
srcset="https://greptile-static-assets.s3.amazonaws.com/badges/FixInCodex.svg?v=6"><img
alt="Fix in Codex"
src="https://greptile-static-assets.s3.amazonaws.com/badges/FixInCodex.svg?v=6"></picture></a>
<!-- /greptile_failed_comments -->

<sub>Reviews (2): Last reviewed commit: ["Normalize local OpenAI v1 base
URLs"](https://github.com/alishahryar1/free-claude-code/commit/2355eac247a6e411f89a781f46e784636ced98d6)
| [Re-trigger
Greptile](https://app.greptile.com/api/retrigger?id=43572841)</sub>

<!-- /greptile_comment -->
2026-07-11 17:53:42 -07:00
Ali Khokhar df038c2470 Retry degraded NVIDIA NIM functions as overloads (#1064)
## Problem

NVIDIA NVCF can return HTTP 400 when a deployed function is DEGRADED,
even though the failure is transient. FCC treated that response as a
deterministic invalid request and never used its provider-owned retry
budget. Fixes #1057.

## Changes

| Before | After |
| --- | --- |
| Every NVIDIA HTTP 400 was non-retryable. | The NVIDIA adapter
recognizes only the structured NVCF `Function id …: DEGRADED function
cannot be invoked` response. |
| A degraded function failed after one attempt. | The shared limiter
applies its existing five-attempt exponential backoff without mutating
the request or adding another retry loop. |
| Exhaustion returned `INVALID_REQUEST / 400`. | Exhaustion returns
canonical `OVERLOADED / 529` while preserving the redacted upstream HTTP
400 detail and request ID. |
| NVIDIA-specific wording could have leaked into shared classification.
| The exact marker stays NVIDIA-owned; shared provider policy owns
canonical semantics, diagnostics, and scheduling, while API wire mapping
remains unchanged. |
| The degraded-function boundary had no deterministic coverage. | Tests
cover recovery, exact exhaustion, near misses, provider isolation, raw
exception preservation, trace metadata, and credential redaction; all
five CI gates pass with 2,214 tests passed and 7 skipped. |
| Package version was 3.5.7. | Package version is 3.5.8 with an updated
lockfile. |

<!-- greptile_comment -->

<details open><summary><h3>Greptile Summary</h3></summary>

This PR teaches NVIDIA NIM degraded-function failures to use the shared
overload retry path. The main changes are:

- Adds a provider-specific failure override hook.
- Maps the exact NVCF degraded-function HTTP 400 marker to canonical
overload semantics.
- Threads the override through shared retry and final failure
classification.
- Adds focused tests for retry, exhaustion, near misses, provider
isolation, redaction, and trace metadata.
- Updates docs, smoke coverage metadata, package version, and lockfile.
</details>


<h3>Confidence Score: 5/5</h3>

This looks safe to merge after considering the narrow NVIDIA error-body
match.

The retry and final classification flow preserve the raw upstream error,
and non-NVIDIA providers keep the existing 400 behavior.

src/free_claude_code/providers/nvidia_nim/client.py should harden the
degraded marker lookup for nested SDK error bodies.

<details><summary><h3><a href="https://www.greptile.com/trex"><img
alt="T-Rex"
src="https://greptile-static-assets.s3.amazonaws.com/trex/trex_green.svg"
height="20" align="absmiddle"></a> T-Rex Logs</h3></summary>

**What T-Rex did**
- T-Rex produced a proof for the posted P1 finding and linked it to the
corresponding review comment.
- A contract-validation proof was captured showing the full command
transcript with timestamps, working directory, commands, pytest verbose
output, and exit codes.
- The focused Nvidia degraded retry test ran and completed with 8 passed
in 2.43s.
- A broader test run reported worker crashes, with 3 failed and 70
passed.
- A serial test for test\_openai\_compat\_5xx\_retry.py timed out after
a hang, exiting at 124.

<a
href="https://app.greptile.com/trex/runs/14112472/artifacts"><picture><source
media="(prefers-color-scheme: dark)"
srcset="https://greptile-static-assets.s3.amazonaws.com/badges/ViewAllArtifactsDark.svg?v=4"><source
media="(prefers-color-scheme: light)"
srcset="https://greptile-static-assets.s3.amazonaws.com/badges/ViewAllArtifacts.svg?v=4"><img
alt="View all artifacts"
src="https://greptile-static-assets.s3.amazonaws.com/badges/ViewAllArtifacts.svg?v=4"></picture></a>

<sub><a href="https://www.greptile.com/trex"><img alt="T-Rex"
src="https://greptile-static-assets.s3.amazonaws.com/trex/trex_green.svg"
height="14" align="absmiddle"></a> Ran code and verified through
T-Rex</sub>
</details>


<details open><summary><h3>Important Files Changed</h3></summary>




| Filename | Overview |
|----------|----------|
| src/free_claude_code/providers/nvidia_nim/client.py | Adds
NVIDIA-specific degraded-function detection; the match is narrow and
only reads a top-level `detail` field. |
| src/free_claude_code/providers/rate_limit.py | Uses provider overrides
for retry qualification while preserving the raw exception after retry
exhaustion. |
| src/free_claude_code/providers/failure_policy.py | Adds the shared
override type and final classification hook, plus a reusable
overloaded-provider failure constructor. |
| src/free_claude_code/providers/transports/openai_chat/transport.py |
Passes provider-specific failure overrides into stream creation retries
and final failure mapping. |
| tests/providers/test_nvidia_nim_degraded_retry.py | Adds coverage for
NVIDIA degraded-function retry success, retry exhaustion, non-matching
400s, provider isolation, redaction, and traces. |

</details>


<!-- greptile_failed_comments -->
<h3>Comments Outside Diff (1)</h3>

1. General comment 

<a href="#"><img alt="P1"
src="https://greptile-static-assets.s3.amazonaws.com/badges/p1.svg?v=9"
align="top"></a> **Adjacent OpenAI-compatible 5xx exhaustion test
crashes workers and hangs serially**

   - **Bug**
- The requested adjacent retry/rate-limit coverage is not stable on
current head. Under xdist, three parameters of
`test_nim_stream_openai_5xx_exhausted_emits_user_message` caused workers
to terminate unexpectedly, producing an exit code 1. A follow-up serial
rerun passed the 500 parameter but timed out after 60 seconds while
running the 502 parameter, indicating the exhausted 5xx path can hang or
terminate the worker rather than completing deterministically.
   - **Cause**
- The executed evidence points to the OpenAI-compatible NIM 5xx
exhaustion path exercised by
`tests/providers/test_openai_compat_5xx_retry.py::test_nim_stream_openai_5xx_exhausted_emits_user_message`.
The exact root cause was not isolated within the validation budget, but
the affected area overlaps changed retry/transport handling files.
   - **Fix**
- Investigate the OpenAI-compatible streaming 5xx exhaustion path for
non-terminating retry/stream cleanup or worker-fatal behavior. Ensure
exhausted 5xx cases raise the expected user-facing error promptly for
all status parameters, then rerun the adjacent command and the targeted
serial parametrized test.

<sub><a href="https://www.greptile.com/trex"><img alt="T-Rex"
src="https://greptile-static-assets.s3.amazonaws.com/trex/trex_green.svg"
height="14" align="absmiddle"></a> Ran code and verified through
T-Rex</sub>

<!-- /greptile_failed_comments -->

<a
href="https://app.greptile.com/api/ide/codex?prompt=IMPORTANT%3A%20Work%20in%20the%20repository%20%22alishahryar1%2Ffree-claude-code%22%20on%20the%20existing%20branch%20%22ali%2Fretry-nim-degraded-function%22.%20Checkout%20that%20branch%20%E2%80%94%20do%20NOT%20create%20a%20new%20branch%20or%20open%20a%20new%20PR.%20Push%20your%20changes%20to%20%22ali%2Fretry-nim-degraded-function%22.%0A%0AFix%20the%20following%201%20code%20review%20issue.%20Work%20through%20them%20one%20at%20a%20time%2C%20proposing%20concise%20fixes.%0A%0A---%0A%0A%23%23%23%20Issue%201%20of%201%0Asrc%2Ffree_claude_code%2Fproviders%2Fnvidia_nim%2Fclient.py%3A126-127%0A**Nested%20Error%20Body%20Misses%20Retry**%0A%0AWhen%20the%20OpenAI%20SDK%20exposes%20a%20400%20response%20as%20the%20common%20nested%20error%20shape%2C%20such%20as%20%60%7B%22error%22%3A%20%7B%22message%22%3A%20%22Function%20id%20...%3A%20DEGRADED%20function%20cannot%20be%20invoked%22%7D%7D%60%2C%20this%20lookup%20never%20sees%20the%20degraded%20marker.%20That%20path%20falls%20back%20to%20the%20shared%20400%20handling%2C%20so%20a%20transient%20degraded%20NVCF%20function%20still%20fails%20after%20one%20attempt%20instead%20of%20using%20the%20provider%20retry%20budget.%0A%0A&repo=alishahryar1%2Ffree-claude-code&pr=1064&platform=github"><picture><source
media="(prefers-color-scheme: dark)"
srcset="https://greptile-static-assets.s3.amazonaws.com/badges/FixAllInCodexDark.svg?v=6"><source
media="(prefers-color-scheme: light)"
srcset="https://greptile-static-assets.s3.amazonaws.com/badges/FixAllInCodex.svg?v=6"><img
alt="Fix All in Codex"
src="https://greptile-static-assets.s3.amazonaws.com/badges/FixAllInCodex.svg?v=6"></picture></a>

<sub>Reviews (1): Last reviewed commit: ["Retry degraded NVIDIA NIM
functions as
o..."](https://github.com/alishahryar1/free-claude-code/commit/2cc7f056a73d1baefb1e1395b181d26efc32dc9e)
| [Re-trigger
Greptile](https://app.greptile.com/api/retrigger?id=43555691)</sub>

> Greptile also left **1 inline comment** on this PR.

<!-- /greptile_comment -->
2026-07-11 14:28:41 -07:00
Ali Khokhar 5ffa47fbc3 Make response stream lifetimes explicit (#1060)
## Problem

Client disconnects and response-start send failures could abandon a
prefetched provider stream and its generation lease. Re-yielding
iterators and response-proxy middleware left no owner that closed the
complete body chain before runtime release.

## Changes

| Before | After |
| --- | --- |
| Starlette body iteration indirectly owned stream cleanup and lease
release. | One FCC streaming response surrounds the real ASGI send,
closes the body transitively, then releases the lease exactly once. |
| The prefetched first-frame generator could not close its tail before
replay began. | An explicit closeable replay iterator owns the
prefetched tail in every commit state. |
| Tracing, execution, Responses conversion, and native transport
transforms re-yielded inputs without closing them. | Every retained
transform closes its direct input; redundant transport wrappers are
removed while provider construction failures remain deferred. |
| Function-style correlation middleware proxied and canceled streaming
responses. | Pure ASGI correlation spans the complete stream, preserves
request headers and log context, and keeps the catch-all 500 fallback
correlated. |
| Repeated cancellation could interrupt pre-start and post-start
cleanup. | Shielded completion tasks finish body closure before release
and then restore caller cancellation. |
| The package version was 3.5.5. | The package version is 3.5.6; full CI
passes with 2,162 tests and stable live API/provider/disconnect/client
smoke passes 63 scenarios. |

<!-- greptile_comment -->

<details open><summary><h3>Greptile Summary</h3></summary>

This PR makes streaming response ownership explicit across the API path.
The main changes are:

- Adds a managed streaming response that closes the body chain before
releasing provider resources.
- Adds a prefetched replay iterator for first-frame commit handling.
- Moves request correlation to pure ASGI middleware for full-stream
context.
- Propagates direct-input closure through execution, tracing, Responses
conversion, and provider transports.
- Bumps the package version and updates tests for stream cleanup
behavior.
</details>

<h3>Confidence Score: 5/5</h3>

This looks safe to merge.

No blocking issues found in the changed code.

None.

<details><summary><h3><a href="https://www.greptile.com/trex"><img
alt="T-Rex"
src="https://greptile-static-assets.s3.amazonaws.com/trex/trex_green.svg"
height="20" align="absmiddle"></a> T-Rex Logs</h3></summary>

**What T-Rex did**
- Validated the execution environment by reviewing the environment proof
log, confirming uv 0.11.28, CPython 3.14.0, a repo-local virtual
environment, and exit code 0.
- Verified that the requested test command was executed, based on the
test proof log.
- Confirmed the test run completed successfully with 91 tests passing in
3.63 seconds, as shown in the test proof log.

<a
href="https://app.greptile.com/trex/runs/14099258/artifacts"><picture><source
media="(prefers-color-scheme: dark)"
srcset="https://greptile-static-assets.s3.amazonaws.com/badges/ViewAllArtifactsDark.svg?v=4"><source
media="(prefers-color-scheme: light)"
srcset="https://greptile-static-assets.s3.amazonaws.com/badges/ViewAllArtifacts.svg?v=4"><img
alt="View all artifacts"
src="https://greptile-static-assets.s3.amazonaws.com/badges/ViewAllArtifacts.svg?v=4"></picture></a>

<sub><a href="https://www.greptile.com/trex"><img alt="T-Rex"
src="https://greptile-static-assets.s3.amazonaws.com/trex/trex_green.svg"
height="14" align="absmiddle"></a> Ran code and verified through
T-Rex</sub>
</details>

<details open><summary><h3>Important Files Changed</h3></summary>

| Filename | Overview |
|----------|----------|
| src/free_claude_code/api/response_streams.py | Adds the managed
response owner, first-frame replay iterator, and shielded cleanup flow.
|
| src/free_claude_code/api/request_ids.py | Adds pure ASGI request
correlation and response-start header injection. |
| src/free_claude_code/core/trace.py | Adds shared stream input closure
tracing and closes traced inputs on exit. |
| src/free_claude_code/application/execution.py | Closes provider stream
iterators from the executor wrapper when streaming ends. |
|
src/free_claude_code/providers/transports/anthropic_messages/transport.py
| Returns provider runner streams directly and closes layered SSE
iterators explicitly. |

</details>

<sub>Reviews (2): Last reviewed commit: ["Make response stream lifetimes
explicit"](https://github.com/alishahryar1/free-claude-code/commit/cb698c62c08924d5f80a1cea7dbd19c0b8af26a2)
| [Re-trigger
Greptile](https://app.greptile.com/api/retrigger?id=43527620)</sub>

<!-- /greptile_comment -->
2026-07-11 09:33:31 -07:00
Ali Khokhar 5b35ac7c49 Make provider stream ownership exact and typed (#1059)
## Problem

Provider stream runners lived in sibling modules, accepted `transport:
Any`, and reached through private transport state. The split hid
ownership behind an import-cycle workaround and prevented type checking
from protecting the collaboration.

## Changes

| Before | After |
| --- | --- |
| OpenAI-chat and native Messages streaming and recovery were split
across six modules. | Each transport module owns one exactly typed
private request runner and its recovery lifecycle; four obsolete modules
are deleted without shims. |
| Native streaming retained an unused mode export and transformation
hooks. | Native streaming directly owns the single normalization path
used by llama.cpp and Ollama. |
| Provider retry, holdback, continuation, tool salvage, tracing, and
subclass hooks crossed an untyped backchannel. | Existing behavior and
provider-specific hooks remain unchanged behind checked
`OpenAIChatTransport` and `AnthropicMessagesTransport` collaborators. |
| Architecture checks did not cover this internal ownership boundary. |
A generic AST contract rejects untyped transport collaborators and
cross-module private transport access. |
| Package version was `3.5.4`. | Package version is `3.5.5`; full CI
passes with 2,136 tests and configured-provider behavior was exercised
live. |

<!-- greptile_comment -->

<details open><summary><h3>Greptile Summary</h3></summary>

This PR moves provider streaming ownership into typed transport modules.
The main changes are:

- OpenAI-chat stream and recovery logic now lives in
`openai_chat/transport.py`.
- Anthropic Messages stream and recovery logic now lives in
`anthropic_messages/transport.py`.
- Obsolete stream and recovery helper modules were removed.
- Import-boundary tests now check typed transport collaborators and
private transport access.
- The package version and architecture notes were updated.
</details>

<h3>Confidence Score: 5/5</h3>

This looks safe to merge.

No blocking issues found in the changed code.

<details><summary><h3><a href="https://www.greptile.com/trex"><img
alt="T-Rex"
src="https://greptile-static-assets.s3.amazonaws.com/trex/trex_green.svg"
height="20" align="absmiddle"></a> T-Rex Logs</h3></summary>

**What T-Rex did**
- Ran the provider contract tests; the test suite completed with 19
passing tests and EXIT\_CODE 0.
- Initiated the provider regression test run, which was terminated
before completion with EXIT\_CODE 137.
- Executed Ruff provider checks; all checks passed with EXIT\_CODE 0.
- Executed Ty provider type checks; all checks passed with EXIT\_CODE 0.

<a
href="https://app.greptile.com/trex/runs/14096487/artifacts"><picture><source
media="(prefers-color-scheme: dark)"
srcset="https://greptile-static-assets.s3.amazonaws.com/badges/ViewAllArtifactsDark.svg?v=4"><source
media="(prefers-color-scheme: light)"
srcset="https://greptile-static-assets.s3.amazonaws.com/badges/ViewAllArtifacts.svg?v=4"><img
alt="View all artifacts"
src="https://greptile-static-assets.s3.amazonaws.com/badges/ViewAllArtifacts.svg?v=4"></picture></a>

<sub><a href="https://www.greptile.com/trex"><img alt="T-Rex"
src="https://greptile-static-assets.s3.amazonaws.com/trex/trex_green.svg"
height="14" align="absmiddle"></a> Ran code and verified through
T-Rex</sub>
</details>

<details open><summary><h3>Important Files Changed</h3></summary>

| Filename | Overview |
|----------|----------|
|
src/free_claude_code/providers/transports/anthropic_messages/transport.py
| Native streaming and recovery were moved into a typed per-request
runner. |
| src/free_claude_code/providers/transports/openai_chat/transport.py |
OpenAI-chat streaming and recovery were moved into a typed per-request
runner. |
| tests/contracts/test_import_boundaries.py | The transport boundary
contract now rejects nested `Any` transport annotations and private
transport access from helper modules. |

</details>

<sub>Reviews (2): Last reviewed commit: ["Make provider stream owners
exact and
ty..."](https://github.com/alishahryar1/free-claude-code/commit/30d0c88e7a28e4c7ac44cc622f52deca5cf83b93)
| [Re-trigger
Greptile](https://app.greptile.com/api/retrigger?id=43519560)</sub>

<!-- /greptile_comment -->
2026-07-11 08:37:49 -07:00
Ali Khokhar 23ed6bc87e Make provider backoff admission monotonic (#1053)
## Problem

Concurrent provider requests could miss or shorten a reactive backoff
while waiting for proactive admission. Separate gate commits could also
waste quota or release requests as an expiry burst.

## Changes

| Before | After |
| --- | --- |
| Reactive deadlines were checked once and replaced by later backoffs. |
Reactive deadlines are rechecked until clear and only extended by later
backoffs. |
| Proactive capacity was recorded before the final reactive decision. |
Conditional commit records proactive capacity only while the reactive
deadline is clear. |
| The limiter exposed an ambiguous `set_blocked` callback. | Both
transport families use the explicit `extend_reactive_block` operation. |
| Cross-gate races and wait cancellation lacked deterministic coverage.
| Deterministic tests cover conditional commit, extensions,
non-shortening, and cancellation. |
| The architecture left gate interaction implicit. | The architecture
defines monotonic two-gate admission without wasted quota or expiry
bursts. |

<!-- greptile_comment -->

<details open><summary><h3>Greptile Summary</h3></summary>

This PR makes provider backoff admission monotonic. The main changes
are:

- Conditional proactive admission before recording quota.
- Reactive waits that recheck extended deadlines.
- Monotonic reactive block extension for retry and stream failure paths.
- Tests for admission retries, deadline extension, non-shortening, and
cancellation.
- Version and architecture updates for the new limiter behavior.
</details>

<h3>Confidence Score: 5/5</h3>

This looks safe to merge.

No blocking issues found in the changed code.

None.

<details><summary><h3><a href="https://www.greptile.com/trex"><img
alt="T-Rex"
src="https://greptile-static-assets.s3.amazonaws.com/trex/trex_green.svg"
height="20" align="absmiddle"></a> T-Rex Logs</h3></summary>

**What T-Rex did**
- Ran the provider-rate-limit pytest suite; all 53 tests completed and
passed in 4.02 seconds with EXIT\_CODE: 0.
- Started the transport-backoff pytest run; it advanced through most
tests before xdist workers were terminated and the process exited with
EXIT\_CODE: 143.
- Verified that no live provider credentials, Docker containers, or
external service smoke tests were used.

<a
href="https://app.greptile.com/trex/runs/14088781/artifacts"><picture><source
media="(prefers-color-scheme: dark)"
srcset="https://greptile-static-assets.s3.amazonaws.com/badges/ViewAllArtifactsDark.svg?v=4"><source
media="(prefers-color-scheme: light)"
srcset="https://greptile-static-assets.s3.amazonaws.com/badges/ViewAllArtifacts.svg?v=4"><img
alt="View all artifacts"
src="https://greptile-static-assets.s3.amazonaws.com/badges/ViewAllArtifacts.svg?v=4"></picture></a>

<sub><a href="https://www.greptile.com/trex"><img alt="T-Rex"
src="https://greptile-static-assets.s3.amazonaws.com/trex/trex_green.svg"
height="14" align="absmiddle"></a> Ran code and verified through
T-Rex</sub>
</details>

<details open><summary><h3>Important Files Changed</h3></summary>

| Filename | Overview |
|----------|----------|
| src/free_claude_code/core/rate_limit.py | Adds conditional
sliding-window admission without recording quota on rejection. |
| src/free_claude_code/providers/rate_limit.py | Reworks provider
admission to combine reactive waits with final proactive commit checks.
|
| src/free_claude_code/providers/transports/anthropic_messages/stream.py
| Routes stream rate-limit marking through monotonic reactive block
extension. |
| src/free_claude_code/providers/transports/openai_chat/stream.py |
Routes stream rate-limit marking through monotonic reactive block
extension. |
| tests/core/test_strict_sliding_window.py | Covers rejected conditional
admission and commit-time timestamp recording. |
| tests/providers/test_provider_rate_limit.py | Covers reactive
admission retries, extended deadlines, cancellation, and non-shortening
behavior. |

</details>

<sub>Reviews (2): Last reviewed commit: ["Make provider backoff
admission
monotoni..."](https://github.com/alishahryar1/free-claude-code/commit/a5c637393d1996e51c195875636ecb0536522932)
| [Re-trigger
Greptile](https://app.greptile.com/api/retrigger?id=43504969)</sub>

<!-- /greptile_comment -->
2026-07-11 05:07:47 -07:00
Ali Khokhar d428b5904a Replace historical architecture tests with declarative boundaries (#1050)
## Problem

Architecture contracts mixed real dependency rules with deleted-module
tombstones and exact internal file inventories. Correct refactors
therefore had to preserve history instead of the current ownership
model.

## Changes

| Before | After |
| --- | --- |
| Cross-package rules were duplicated across narrow source and layout
assertions. | One least-privilege matrix and AST scanner enforce every
production package edge. |
| Bare imports, undeclared ownership roots, and module cycles could
escape generic enforcement. | Namespaced imports, initialized owners,
exact exceptions, and an acyclic module graph are enforced. |
| Responses, messaging-tree, and optional dependency ownership relied on
scattered checks. | Facade use and lazy optional dependency owners are
declared and verified centrally. |
| The messaging facade re-exported workflow, persistence, and parsing
internals. | The messaging facade exposes only ingress values, platform
ports, and managed-session protocols. |
| Migration-era tests froze deleted modules and internal filenames. |
Customer contracts remain while obsolete tombstones and layout
inventories are removed. |

<!-- greptile_comment -->

<details open><summary><h3>Greptile Summary</h3></summary>

This PR replaces historical architecture checks with declarative
import-boundary rules. The main changes are:

- Added a documented package dependency matrix and facade ownership
rules.
- Narrowed the messaging package facade to its supported extension
surface.
- Moved messaging tree consumers to the `messaging.trees` facade.
- Reworked architecture tests around AST scanning, optional dependency
owners, and acyclic imports.
- Bumped the package version and refreshed the lockfile.
</details>

<h3>Confidence Score: 5/5</h3>

This looks safe to merge.

No blocking issues found in the changed code.

<details><summary><h3><a href="https://www.greptile.com/trex"><img
alt="T-Rex"
src="https://greptile-static-assets.s3.amazonaws.com/trex/trex_green.svg"
height="20" align="absmiddle"></a> T-Rex Logs</h3></summary>

**What T-Rex did**
- The architecture contracts tests were executed as part of the general
contract validation.
- The test run completed with exit code 0, indicating success.
- The pytest summary shows 26 tests passed in 3.99 seconds.
- An artifact log captures the exact command, working directory,
environment recreation output, and verbose test item names for audit.

<a
href="https://app.greptile.com/trex/runs/14085499/artifacts"><picture><source
media="(prefers-color-scheme: dark)"
srcset="https://greptile-static-assets.s3.amazonaws.com/badges/ViewAllArtifactsDark.svg?v=4"><source
media="(prefers-color-scheme: light)"
srcset="https://greptile-static-assets.s3.amazonaws.com/badges/ViewAllArtifacts.svg?v=4"><img
alt="View all artifacts"
src="https://greptile-static-assets.s3.amazonaws.com/badges/ViewAllArtifacts.svg?v=4"></picture></a>

<sub><a href="https://www.greptile.com/trex"><img alt="T-Rex"
src="https://greptile-static-assets.s3.amazonaws.com/trex/trex_green.svg"
height="14" align="absmiddle"></a> Ran code and verified through
T-Rex</sub>
</details>

<details open><summary><h3>Important Files Changed</h3></summary>

| Filename | Overview |
|----------|----------|
| src/free_claude_code/messaging/__init__.py | Narrows the top-level
messaging exports to the documented supported extension types. |
| tests/contracts/test_import_boundaries.py | Replaces historical layout
checks with declarative import-boundary enforcement. |
| ARCHITECTURE.md | Documents the current dependency matrix, facade
boundaries, and optional dependency owners. |
| src/free_claude_code/messaging/node_event_pipeline.py | Uses the
messaging tree facade for `NodeClaim`. |
| src/free_claude_code/messaging/node_runner.py | Uses the messaging
tree facade for queue, snapshot, cancellation, and claim types. |

</details>

<sub>Reviews (2): Last reviewed commit: ["Enforce architecture with
declarative
im..."](https://github.com/alishahryar1/free-claude-code/commit/68508f3d2bad19d4d1b8a9ca5d181e1433db4d71)
| [Re-trigger
Greptile](https://app.greptile.com/api/retrigger?id=43498975)</sub>

<!-- /greptile_comment -->
2026-07-11 03:25:01 -07:00
Ali Khokhar 701a394697 Make provider capabilities semantic and typed (#1047)
## Problem

Provider metadata mixed descriptive strings and transport-family names
with product policy. The Messages API had to know provider
implementation details to decide server-tool behavior, while locality
and thinking metadata had ambiguous owners.

## Changes

| Before | After |
| --- | --- |
| Provider descriptors exposed transport names and untyped capability
strings. | Provider descriptors expose an immutable semantic value
containing only locality and server-tool passthrough. |
| The Messages handler derived provider sets from transport families. |
`ResolvedModel` carries catalog capabilities and the handler asks the
resolved semantic policy. |
| Discovery, Admin, and smoke selection inferred locality from
credentials or string membership. | Discovery, Admin, and smoke
selection read the catalog-owned `local` capability. |
| Provider-wide thinking strings competed with discovered model
metadata. | `ProviderModelInfo.supports_thinking` remains the sole
per-model thinking authority. |
| Invalid mapped providers could become raw catalog lookup failures
during the migration. | Routing and provider construction share the
canonical typed unknown-provider error. |
| LM Studio smoke contracts described a native Anthropic transport. | LM
Studio smoke contracts describe its OpenAI-chat-backed Messages path. |
| Tests encoded transport identities as product behavior. | Tests prove
both directions of capability-versus-identity independence while
preserving existing wire behavior and diagnostics. |

<!-- greptile_comment -->

<details open><summary><h3>Greptile Summary</h3></summary>

This PR moves provider policy from transport names to typed catalog
capabilities. The main changes are:

- Added immutable provider capability values for locality and
server-tool passthrough.
- Routed provider capabilities through `ResolvedModel` for Messages API
policy.
- Updated admin status, discovery, smoke config, docs, and tests to read
semantic capabilities.
- Centralized unknown-provider diagnostics for routing and provider
construction.
- Bumped the package version and lockfile to `3.4.22`.
</details>

<h3>Confidence Score: 5/5</h3>

This looks safe to merge.

No blocking issues found in the changed code.

None.

<details><summary><h3><a href="https://www.greptile.com/trex"><img
alt="T-Rex"
src="https://greptile-static-assets.s3.amazonaws.com/trex/trex_green.svg"
height="20" align="absmiddle"></a> T-Rex Logs</h3></summary>

**What T-Rex did**
- The Pytest run for the provider capabilities changes was executed and
completed successfully, as shown by the log's EXIT\_CODE: 0.
- The Ruff linter run for the provider capabilities changes was executed
and completed successfully, as shown by the log's EXIT\_CODE: 0.
- The focused suite exercised the changed code paths via repository
tests rather than relying solely on static inspection.

<a
href="https://app.greptile.com/trex/runs/14074560/artifacts"><picture><source
media="(prefers-color-scheme: dark)"
srcset="https://greptile-static-assets.s3.amazonaws.com/badges/ViewAllArtifactsDark.svg?v=4"><source
media="(prefers-color-scheme: light)"
srcset="https://greptile-static-assets.s3.amazonaws.com/badges/ViewAllArtifacts.svg?v=4"><img
alt="View all artifacts"
src="https://greptile-static-assets.s3.amazonaws.com/badges/ViewAllArtifacts.svg?v=4"></picture></a>

<sub><a href="https://www.greptile.com/trex"><img alt="T-Rex"
src="https://greptile-static-assets.s3.amazonaws.com/trex/trex_green.svg"
height="14" align="absmiddle"></a> Ran code and verified through
T-Rex</sub>
</details>

<details open><summary><h3>Important Files Changed</h3></summary>

| Filename | Overview |
|----------|----------|
| src/free_claude_code/config/provider_catalog.py | Replaces transport
and string capability metadata with typed immutable provider
capabilities. |
| src/free_claude_code/application/routing.py | Adds catalog
capabilities to resolved models and uses the shared unknown-provider
error. |
| src/free_claude_code/api/handlers/messages.py | Uses resolved semantic
capabilities to decide server-tool passthrough. |
| src/free_claude_code/config/admin/status.py | Classifies local
providers through the new catalog capability. |
| src/free_claude_code/providers/runtime/discovery.py | Uses typed
locality to decide which providers are discovered only when referenced.
|
| src/free_claude_code/application/errors.py | Adds a shared constructor
for canonical unknown-provider messages. |

</details>

<sub>Reviews (1): Last reviewed commit: ["Make provider capabilities
semantic and
..."](https://github.com/alishahryar1/free-claude-code/commit/b2732ef74d61c632ccae332a19aaa655f45a96e7)
| [Re-trigger
Greptile](https://app.greptile.com/api/retrigger?id=43475337)</sub>

<!-- /greptile_comment -->
2026-07-10 22:48:47 -07:00
Ali Khokhar e22a38b2c2 Canonicalize provider failure and retry ownership (#1046)
## Problem

Provider SDK classification, retry policy, canonical failures, and
downstream wire errors shared exception types across layers. That
blurred ownership and let cleanup or provisional Responses tool failures
mask the real provider diagnostic.

## Changes

| Before | After |
| --- | --- |
| Provider failures carried Anthropic wire types and core code
classified OpenAI/httpx errors. | Protocol-neutral `ExecutionFailure`
values cross layers, providers classify SDK errors, and protocol
packages map wire types. |
| Provider adapters could author terminal wire events. | The HTTP commit
boundary selects non-2xx JSON or a protocol terminal event with one
ingress request ID. |
| Retry policy and diagnostic handling were spread across core and
provider modules. | Providers own the unchanged retry budgets while
neutral core utilities own bounded credential redaction. |
| Stream cleanup could replace an already-mapped provider failure. |
Cleanup records safe metadata and preserves the canonical failure,
status, and diagnostic. |
| An incomplete Responses tool could preempt a later provider failure. |
Tool-finalization errors remain provisional so canonical provider
failures take precedence. |
| Readiness failures reused provider exception types. |
Application-owned errors represent deterministic validation and
availability phases without terminal retry headers. |
| Legacy exception and recovery owners remained importable. | Obsolete
modules are deleted without shims, architecture rules enforce the
boundaries, and package version is 3.4.21. |

<!-- greptile_comment -->

<details open><summary><h3>Greptile Summary</h3></summary>

This PR canonicalizes provider failure handling across the API boundary.
The main changes are:

- Adds protocol-neutral execution failure values and safe diagnostics.
- Moves SDK and HTTP failure classification into provider-owned policy.
- Lets Messages and Responses choose their own wire error payloads.
- Preserves canonical failures across stream cleanup and committed
stream failures.
- Makes incomplete Responses tool errors provisional until finalization.
</details>

<h3>Confidence Score: 5/5</h3>

This looks safe to merge.

No blocking issues found in the changed code.

No files need attention.

<details><summary><h3><a href="https://www.greptile.com/trex"><img
alt="T-Rex"
src="https://greptile-static-assets.s3.amazonaws.com/trex/trex_green.svg"
height="20" align="absmiddle"></a> T-Rex Logs</h3></summary>

**What T-Rex did**
- Ran the API failure contract suite and related tests
(tests/api/test\_execution\_failure\_contract.py,
tests/core/test\_failure\_protocol\_mapping.py,
tests/providers/test\_execution\_failure\_boundary.py,
tests/providers/test\_failure\_policy.py); 48 passed in 3.36s.
- Ran the streaming boundaries tests including response streams, stream
recovery, and streaming errors; 70 passed in 5.36s.
- Ran the OpenAI responses tests; 20 passed in 4.42s.

<a
href="https://app.greptile.com/trex/runs/14071383/artifacts"><picture><source
media="(prefers-color-scheme: dark)"
srcset="https://greptile-static-assets.s3.amazonaws.com/badges/ViewAllArtifactsDark.svg?v=4"><source
media="(prefers-color-scheme: light)"
srcset="https://greptile-static-assets.s3.amazonaws.com/badges/ViewAllArtifacts.svg?v=4"><img
alt="View all artifacts"
src="https://greptile-static-assets.s3.amazonaws.com/badges/ViewAllArtifacts.svg?v=4"></picture></a>

<sub><a href="https://www.greptile.com/trex"><img alt="T-Rex"
src="https://greptile-static-assets.s3.amazonaws.com/trex/trex_green.svg"
height="14" align="absmiddle"></a> Ran code and verified through
T-Rex</sub>
</details>

<details open><summary><h3>Important Files Changed</h3></summary>

| Filename | Overview |
|----------|----------|
| src/free_claude_code/providers/transports/http.py | Adds cleanup-safe
stream closing that preserves established outcomes. |
| src/free_claude_code/core/openai_responses/stream.py | Preserves
canonical execution failures when committed Responses streams fail. |
| src/free_claude_code/core/openai_responses/streaming/assembler.py |
Keeps malformed tool-call errors provisional so later provider failures
can win. |
| src/free_claude_code/core/failures.py | Defines neutral failure kinds
and exception-group lookup for execution failures. |

</details>

<sub>Reviews (2): Last reviewed commit: ["Preserve canonical outcomes in
grouped
a..."](https://github.com/alishahryar1/free-claude-code/commit/f57f21241dbe582985627ed4fb40734b2c656809)
| [Re-trigger
Greptile](https://app.greptile.com/api/retrigger?id=43468297)</sub>

<!-- /greptile_comment -->
2026-07-10 21:33:50 -07:00
Ali Khokhar f8c21a48f2 Introduce a typed application boundary for provider execution (#1045)
## Problem

The HTTP adapter owned model routing, provider execution, and
runtime-facing contracts, so API handlers depended on provider
implementation types. Provider preflight also discovered private request
builders dynamically, obscuring the boundary that must fail before
streaming begins.

## Changes

| Before | After |
| --- | --- |
| `api/` owned model routing and shared provider execution. |
`application/` owns routing and a settings-independent
`ProviderExecutor`. |
| API handlers accepted `BaseProvider` callbacks. | API handlers consume
the narrow structural `ProviderPort`. |
| `BaseProvider` discovered `_build_request_body` dynamically. | Both
transport families implement explicit abstract preflight, with LM Studio
composing context validation. |
| Request leases, task control, and provider model metadata had
adapter/provider owners. | Application-owned ports and immutable values
define those cross-package contracts. |
| Boundary direction was implicit. | Architecture contracts and
documentation enforce the final dependency direction. |
| Package version was `3.4.19`. | Package version is `3.4.20`, with the
lockfile updated. |
| Coverage followed the old module layout. | Deterministic
boundary/preflight regressions and live Messages/Responses smokes cover
the new shape. |

<!-- greptile_comment -->

<details open><summary><h3>Greptile Summary</h3></summary>

This PR adds a typed application boundary for provider execution. The
main changes are:

- New `application` package for routing, execution, ports, and model
metadata.
- API handlers now call application-owned routing and provider
execution.
- Provider preflight is now explicit on the transport families.
- Runtime API composition now uses a task-control port for `/stop`.
- Import-boundary tests, smoke references, docs, version, and lockfile
were updated.
</details>

<h3>Confidence Score: 5/5</h3>

This looks safe to merge.

No blocking issues found in the changed code.

None.

<details><summary><h3><a href="https://www.greptile.com/trex"><img
alt="T-Rex"
src="https://greptile-static-assets.s3.amazonaws.com/trex/trex_green.svg"
height="20" align="absmiddle"></a> T-Rex Logs</h3></summary>

**What T-Rex did**
- Ran a deterministic pytest run for the provider boundary preflight,
which completed with exit code 0 and 102 tests passing in 4.77 seconds.
- Launched the environment presence check as part of the preflight,
which completed with exit code 0 and confirmed
OPENCODE\_API\_KEY=\[REDACTED\] matched.
- Attempted the live provider smoke test, which completed with exit code
0 and 2 tests skipped due to incomplete smoke configuration.

<a
href="https://app.greptile.com/trex/runs/14067905/artifacts"><picture><source
media="(prefers-color-scheme: dark)"
srcset="https://greptile-static-assets.s3.amazonaws.com/badges/ViewAllArtifactsDark.svg?v=4"><source
media="(prefers-color-scheme: light)"
srcset="https://greptile-static-assets.s3.amazonaws.com/badges/ViewAllArtifacts.svg?v=4"><img
alt="View all artifacts"
src="https://greptile-static-assets.s3.amazonaws.com/badges/ViewAllArtifacts.svg?v=4"></picture></a>

<sub><a href="https://www.greptile.com/trex"><img alt="T-Rex"
src="https://greptile-static-assets.s3.amazonaws.com/trex/trex_green.svg"
height="14" align="absmiddle"></a> Ran code and verified through
T-Rex</sub>
</details>

<details open><summary><h3>Important Files Changed</h3></summary>

| Filename | Overview |
|----------|----------|
| src/free_claude_code/application/execution.py | Moves shared provider
execution into the application layer and keeps eager preflight before
token counting and streaming. |
| src/free_claude_code/application/ports.py | Adds structural provider,
request-runtime, and task-control protocols used across the new
boundary. |
| src/free_claude_code/api/routes.py | Updates route composition to use
the application provider resolver and task-control stop path. |
| src/free_claude_code/providers/base.py | Makes provider preflight
explicit by requiring subclasses or transport bases to implement it. |
| src/free_claude_code/providers/transports/openai_chat/transport.py |
Adds OpenAI-chat preflight through the same request-body builder used by
streaming. |
|
src/free_claude_code/providers/transports/anthropic_messages/transport.py
| Adds native Messages preflight through the native request-body
builder. |
| src/free_claude_code/providers/model_listing.py | Keeps provider
model-list parsing while moving `ProviderModelInfo` ownership to the
application layer. |
| src/free_claude_code/runtime/bootstrap.py | Passes the runtime object
through the new `tasks` service slot. |
| tests/contracts/test_import_boundaries.py | Extends import-boundary
tests for the new application package. |

</details>

<details open><summary><h3>Flowchart</h3></summary>

<a href="#gh-light-mode-only">

```mermaid
%%{init: {'theme': 'neutral'}}%%
flowchart LR
  API[api handlers and routes] --> Routing[application.routing]
  API --> Executor[application.execution]
  API --> Ports[application.ports]
  Executor --> ProviderPort[ProviderPort]
  ProviderPort --> Preflight[preflight_stream]
  ProviderPort --> Stream[stream_response]
  Runtime[runtime bootstrap and provider manager] --> Ports
  Providers[providers] --> Metadata[application.model_metadata]
  Executor --> Core[core anthropic and trace]
  Routing --> Config[config settings and model refs]
```

</a>
<a href="#gh-dark-mode-only">

```mermaid
%%{init: {'theme': 'base', 'themeVariables': {"darkMode": true, "background": "#0d1117", "primaryColor": "#21262d", "primaryTextColor": "#e6edf3", "primaryBorderColor": "#8b949e", "lineColor": "#8b949e", "textColor": "#e6edf3", "edgeLabelBackground": "#161b22", "actorBkg": "#21262d", "actorBorder": "#8b949e", "actorTextColor": "#e6edf3", "actorLineColor": "#8b949e", "signalColor": "#8b949e", "signalTextColor": "#e6edf3", "noteBkgColor": "#373320", "noteBorderColor": "#d4a72c", "noteTextColor": "#f0e6c0", "labelBoxBkgColor": "#21262d", "labelBoxBorderColor": "#8b949e", "labelTextColor": "#e6edf3", "loopTextColor": "#e6edf3", "activationBkgColor": "#30363d", "activationBorderColor": "#8b949e"}}}%%
flowchart LR
  API[api handlers and routes] --> Routing[application.routing]
  API --> Executor[application.execution]
  API --> Ports[application.ports]
  Executor --> ProviderPort[ProviderPort]
  ProviderPort --> Preflight[preflight_stream]
  ProviderPort --> Stream[stream_response]
  Runtime[runtime bootstrap and provider manager] --> Ports
  Providers[providers] --> Metadata[application.model_metadata]
  Executor --> Core[core anthropic and trace]
  Routing --> Config[config settings and model refs]
```

</a>
</details>

<sub>Reviews (1): Last reviewed commit: ["Introduce typed application
boundary"](https://github.com/alishahryar1/free-claude-code/commit/4cdcf97231c812c6f568ca3d74af7ce759d7f2dc)
| [Re-trigger
Greptile](https://app.greptile.com/api/retrigger?id=43462788)</sub>

<!-- /greptile_comment -->
2026-07-10 20:06:57 -07:00
Ali Khokhar 4a0a0360de Move protocol models to their protocol owners (#1044)
## Problem

Anthropic Messages and OpenAI Responses wire models lived under the
inbound API adapter. Neutral protocol and provider code therefore
duck-typed requests, obscuring ownership and weakening dependency
boundaries.

## Changes

| Before | After |
| --- | --- |
| The API package owned Anthropic and Responses protocol models. | Each
protocol package owns and publicly exports its wire models. |
| Core and provider request paths accepted `Any` and probed known fields
with `getattr()`. | Core, transports, and providers consume concrete
`MessagesRequest` values. |
| Responses conversion and streaming received a dumped request mapping.
| Responses conversion and streaming receive one concrete
`OpenAIResponsesRequest`. |
| Anthropic request snapshots lived in generic tracing code. | Anthropic
request snapshots live with the protocol while generic tracing stays
protocol-independent. |
| Protocol tests and provider request doubles reflected the old API
ownership. | Protocol tests live under core and provider tests construct
real wire requests. |
| The API model package mixed protocol and model-catalog schemas. | The
API model package is removed, with catalog schemas beside catalog
construction and no compatibility shim. |
| Package version was `3.4.18`. | Package version is `3.4.19` with an
updated lockfile. |

<!-- greptile_comment -->

<details open><summary><h3>Greptile Summary</h3></summary>

This PR moves protocol request models to their protocol-owned packages.
The main changes are:

- Anthropic Messages models now live under `core.anthropic`.
- OpenAI Responses models now live under `core.openai_responses`.
- API handlers, routes, providers, and tests now use concrete protocol
request types.
- Anthropic request snapshots moved beside the Anthropic protocol
models.
- API model catalog schemas were kept with catalog response
construction.
- The package version and lockfile were updated.
</details>

<h3>Confidence Score: 5/5</h3>

This looks safe to merge.

No blocking issues were found in the changed code. Internal callers were
updated to pass the new concrete protocol models, and no stale internal
imports from the removed API model package were identified.

No files need attention.

<details><summary><h3><a href="https://www.greptile.com/trex"><img
alt="T-Rex"
src="https://greptile-static-assets.s3.amazonaws.com/trex/trex_green.svg"
height="20" align="absmiddle"></a> T-Rex Logs</h3></summary>

**What T-Rex did**
- The Pytest run for protocol ownership focused tests completed, showing
71 passed in 6.24s and EXIT\_CODE: 0.
- A protocol import smoke script was generated for the
import/conversion/trace workflow.
- The protocol import smoke run completed successfully, including model
ownership output, adapter payload evidence, and trace snapshot evidence,
with EXIT\_CODE: 0.

<a
href="https://app.greptile.com/trex/runs/14066157/artifacts"><picture><source
media="(prefers-color-scheme: dark)"
srcset="https://greptile-static-assets.s3.amazonaws.com/badges/ViewAllArtifactsDark.svg?v=4"><source
media="(prefers-color-scheme: light)"
srcset="https://greptile-static-assets.s3.amazonaws.com/badges/ViewAllArtifacts.svg?v=4"><img
alt="View all artifacts"
src="https://greptile-static-assets.s3.amazonaws.com/badges/ViewAllArtifacts.svg?v=4"></picture></a>

<sub><a href="https://www.greptile.com/trex"><img alt="T-Rex"
src="https://greptile-static-assets.s3.amazonaws.com/trex/trex_green.svg"
height="14" align="absmiddle"></a> Ran code and verified through
T-Rex</sub>
</details>

<details open><summary><h3>Important Files Changed</h3></summary>

| Filename | Overview |
|----------|----------|
| src/free_claude_code/core/anthropic/models.py | Anthropic wire request
and response models moved under the Anthropic protocol package. |
| src/free_claude_code/core/anthropic/native_messages_request.py |
Native Anthropic serialization now expects concrete `MessagesRequest`
instances. |
| src/free_claude_code/core/anthropic/conversion.py | OpenAI chat
conversion now reads fields directly from `MessagesRequest`. |
| src/free_claude_code/core/anthropic/request_snapshot.py | Anthropic
request snapshotting moved from generic tracing into the protocol
package. |
| src/free_claude_code/core/openai_responses/models.py | OpenAI
Responses ingress models moved under the Responses protocol package. |
| src/free_claude_code/core/openai_responses/input.py | Responses
conversion now consumes the concrete request model instead of a dumped
mapping. |
| src/free_claude_code/core/openai_responses/streaming/assembler.py |
Responses stream assembly now reads request attributes from
`OpenAIResponsesRequest`. |
| src/free_claude_code/api/routes.py | Routes now import protocol
request models from their new core owners. |
| src/free_claude_code/api/model_catalog.py | Model-list response
schemas now live with model catalog construction. |

</details>

<details open><summary><h3>Flowchart</h3></summary>

<a href="#gh-light-mode-only">

```mermaid
%%{init: {'theme': 'neutral'}}%%
flowchart LR
  API[API routes and handlers] --> Anthropic[core.anthropic models and helpers]
  API --> Responses[core.openai_responses models and adapter]
  Responses --> Anthropic
  Providers[Provider clients and transports] --> Anthropic
  Anthropic --> Trace[core.trace sanitization]
  API --> Catalog[api.model_catalog response schemas]
```

</a>
<a href="#gh-dark-mode-only">

```mermaid
%%{init: {'theme': 'base', 'themeVariables': {"darkMode": true, "background": "#0d1117", "primaryColor": "#21262d", "primaryTextColor": "#e6edf3", "primaryBorderColor": "#8b949e", "lineColor": "#8b949e", "textColor": "#e6edf3", "edgeLabelBackground": "#161b22", "actorBkg": "#21262d", "actorBorder": "#8b949e", "actorTextColor": "#e6edf3", "actorLineColor": "#8b949e", "signalColor": "#8b949e", "signalTextColor": "#e6edf3", "noteBkgColor": "#373320", "noteBorderColor": "#d4a72c", "noteTextColor": "#f0e6c0", "labelBoxBkgColor": "#21262d", "labelBoxBorderColor": "#8b949e", "labelTextColor": "#e6edf3", "loopTextColor": "#e6edf3", "activationBkgColor": "#30363d", "activationBorderColor": "#8b949e"}}}%%
flowchart LR
  API[API routes and handlers] --> Anthropic[core.anthropic models and helpers]
  API --> Responses[core.openai_responses models and adapter]
  Responses --> Anthropic
  Providers[Provider clients and transports] --> Anthropic
  Anthropic --> Trace[core.trace sanitization]
  API --> Catalog[api.model_catalog response schemas]
```

</a>
</details>

<sub>Reviews (1): Last reviewed commit: ["Move protocol models to their
protocol
o..."](https://github.com/alishahryar1/free-claude-code/commit/f1be5c1af4a10da710f80b5f9e7f6044a601f6af)
| [Re-trigger
Greptile](https://app.greptile.com/api/retrigger?id=43459428)</sub>

<!-- /greptile_comment -->
2026-07-10 19:26:27 -07:00
Ali Khokhar 4951983b5e Replace global runtime resources with explicit ownership (#1042)
## Problem

Provider, messaging, and transcription resources relied on
process-global state, leaving replacement, cancellation, and shutdown
ownership ambiguous. Separate server lifetimes could share
event-loop-bound resources or retain failed cleanup work.

## Changes

| Before | After |
| --- | --- |
| Provider clients found limiters through global singleton and scoped
registries. | Each provider instance receives and owns one explicitly
constructed limiter. |
| Messaging queues and voice pipelines relied on singleton or
module-global state. | Each platform owns its limiter and outbox, while
the application owns one injected transcriber. |
| Messaging shutdown mixed ingress, active work, delivery, and SDK
cleanup. | Application shutdown quiesces ingress, drains work, closes
delivery, then releases transcription and providers. |
| Cancelled or failed provider cleanup could be forgotten or treated as
complete. | The provider manager retains shielded generation and
unpublished-runtime cleanup until it succeeds. |
| Discord and Telegram startup tasks could outlive or poison runtime
readiness. | Platform runtimes observe long-lived tasks and retry only
independently repeatable lifecycle steps. |
| Constructor-captured security and diagnostic settings appeared
hot-applicable. | Admin marks those settings restart-required so applied
policy matches the running resource graph. |
| Lifecycle races lacked direct ownership coverage. | Deterministic
cancellation, retry, isolation, teardown, and live smoke contracts
protect the final ownership model. |

<!-- greptile_comment -->

<details open><summary><h3>Greptile Summary</h3></summary>

This PR moves runtime resources from global state into explicitly owned
application objects. The main changes are:

- Provider generations own their rate limiters and cleanup tasks.
- Messaging platforms own their limiter, outbox, ingress, and delivery
lifecycle.
- Application shutdown now runs through ordered cleanup gates.
- Voice transcription is injected as an owned runtime resource.
- Admin config marks constructor-captured settings as restart-required.
</details>

<h3>Confidence Score: 4/5</h3>

The shutdown path needs a bounded cleanup result before merging.

Cleanup steps that hang never reach the retryable incomplete-shutdown
path. ASGI shutdown can remain stuck while waiting for an external SDK,
transcriber, workflow, or provider cleanup. The retry ownership model
works only after cleanup returns or raises.

src/free_claude_code/runtime/application.py

<details><summary><h3><a href="https://www.greptile.com/trex"><img
alt="T-Rex"
src="https://greptile-static-assets.s3.amazonaws.com/trex/trex_green.svg"
height="20" align="absmiddle"></a> T-Rex Logs</h3></summary>

**What T-Rex did**
- T-Rex ran the requested verification, but its local artifact
references were not uploaded.
- The validation run completed successfully with EXIT\_CODE: 0 and 62
tests passed in 3.91 seconds, using the command uv run pytest -vv
tests/runtime/test\_application\_runtime.py
tests/runtime/test\_provider\_manager.py
tests/providers/test\_provider\_runtime.py.

<a
href="https://app.greptile.com/trex/runs/14064214/artifacts"><picture><source
media="(prefers-color-scheme: dark)"
srcset="https://greptile-static-assets.s3.amazonaws.com/badges/ViewAllArtifactsDark.svg?v=4"><source
media="(prefers-color-scheme: light)"
srcset="https://greptile-static-assets.s3.amazonaws.com/badges/ViewAllArtifacts.svg?v=4"><img
alt="View all artifacts"
src="https://greptile-static-assets.s3.amazonaws.com/badges/ViewAllArtifacts.svg?v=4"></picture></a>

<sub><a href="https://www.greptile.com/trex"><img alt="T-Rex"
src="https://greptile-static-assets.s3.amazonaws.com/trex/trex_green.svg"
height="14" align="absmiddle"></a> Ran code and verified through
T-Rex</sub>
</details>

<details open><summary><h3>Important Files Changed</h3></summary>

| Filename | Overview |
|----------|----------|
| src/free_claude_code/runtime/application.py | Refactors shutdown into
ordered retryable cleanup gates, but cleanup awaitables can still block
shutdown forever. |
| src/free_claude_code/runtime/asgi.py | Reports incomplete runtime
shutdown when `close()` returns false. |
| src/free_claude_code/runtime/provider_manager.py | Adds owned provider
cleanup retry state and shielded generation cleanup. |

</details>

<a
href="https://app.greptile.com/api/ide/codex?prompt=IMPORTANT%3A%20Work%20in%20the%20repository%20%22alishahryar1%2Ffree-claude-code%22%20on%20the%20existing%20branch%20%22refactor%2Fruntime-owned-resources%22.%20Checkout%20that%20branch%20%E2%80%94%20do%20NOT%20create%20a%20new%20branch%20or%20open%20a%20new%20PR.%20Push%20your%20changes%20to%20%22refactor%2Fruntime-owned-resources%22.%0A%0AFix%20the%20following%201%20code%20review%20issue.%20Work%20through%20them%20one%20at%20a%20time%2C%20proposing%20concise%20fixes.%0A%0A---%0A%0A%23%23%23%20Issue%201%20of%201%0Asrc%2Ffree_claude_code%2Fruntime%2Fapplication.py%3A59%0A**Cleanup%20Await%20Blocks%20Shutdown**%0A%0AWhen%20a%20platform%20SDK%20stop%2C%20workflow%20drain%2C%20transcriber%20close%2C%20or%20provider%20cleanup%20hangs%2C%20this%20helper%20waits%20forever%20and%20never%20returns%20%60False%60.%20ASGI%20shutdown%20stays%20stuck%20in%20%60runtime.close%28%29%60%20instead%20of%20reporting%20an%20incomplete%20shutdown%2C%20so%20the%20retained%20resource%20graph%20cannot%20be%20retried%20cleanly.%0A%0A&repo=alishahryar1%2Ffree-claude-code&pr=1042&platform=github"><picture><source
media="(prefers-color-scheme: dark)"
srcset="https://greptile-static-assets.s3.amazonaws.com/badges/FixAllInCodexDark.svg?v=6"><source
media="(prefers-color-scheme: light)"
srcset="https://greptile-static-assets.s3.amazonaws.com/badges/FixAllInCodex.svg?v=6"><img
alt="Fix All in Codex"
src="https://greptile-static-assets.s3.amazonaws.com/badges/FixAllInCodex.svg?v=6"></picture></a>

<sub>Reviews (2): Last reviewed commit: ["Report incomplete runtime
shutdown to
AS..."](https://github.com/alishahryar1/free-claude-code/commit/338b2bd179c3875b15bbd52818dd04c780e5d46d)
| [Re-trigger
Greptile](https://app.greptile.com/api/retrigger?id=43454593)</sub>

> Greptile also left **1 inline comment** on this PR.

**Context used:**

- Context used - CLAUDE.md
([source](https://app.greptile.com/alishahryar1/github/Alishahryar1/free-claude-code/-/custom-context?memory=d2fd24d8-0dec-4faf-8ee4-e085e215a2f8))

<!-- /greptile_comment -->
2026-07-10 18:46:50 -07:00
Ali Khokhar 160d63370b Establish single-owner runtime with stream-safe provider hot swaps (#1036)
## Problem

Provider runtime ownership was split between lifecycle code and mutable
FastAPI state, so Admin replacements could leak the new runtime,
double-close the old runtime, or close providers still serving active
streams. The API package also owned concrete process composition,
obscuring subsystem boundaries.

## Changes

| Before | After |
| --- | --- |
| FastAPI routes inspected several concrete `app.state` resources. |
FastAPI receives one explicit `ApiServices` boundary and stores only
`app.state.services`. |
| Admin Apply persisted config and directly replaced one runtime
reference. | Admin Apply validates a candidate, commits atomically, and
publishes it through the single runtime owner. |
| Provider replacement could close clients used by active streams. |
Generation leases retain old providers until each streaming or
non-streaming response finishes. |
| Provider generations owned discovery state and model metadata. |
`ProviderRuntimeManager` owns one application-lifetime catalog and one
discovery task across replacements. |
| API modules composed provider, messaging, and managed CLI resources. |
`runtime.bootstrap` composes concrete subsystems and
`ApplicationRuntime` owns their lifecycle. |
| Admin config, server URLs, and gateway model IDs lived under the API
package. | Admin config lives under `config`, server URLs live under
`config`, and gateway IDs live under `core`. |
| Messaging restoration and shutdown persistence were coordinated
externally. | `MessagingWorkflow` owns snapshot restoration and final
persistence flushing. |
| Hot-swap behavior lacked a real process-level race scenario. |
Deterministic ownership tests and a credential-free subprocess smoke
hold provider A while new requests switch to provider B. |
| Package version was `3.4.16`. | Package version is `3.4.17` with an
updated lockfile. |

<!-- greptile_comment -->

<details open><summary><h3>Greptile Summary</h3></summary>

This PR centralizes server runtime ownership and provider hot swaps. The
main changes are:

- Adds `ApplicationRuntime` and `ProviderRuntimeManager` as the process
owners.
- Moves FastAPI to an explicit `ApiServices` boundary.
- Retains provider generations until request and stream responses
finish.
- Moves admin config, server URL, and gateway model ID modules to
neutral package owners.
- Updates admin apply to validate, persist, and publish provider-only
changes through the runtime owner.
- Adds runtime ownership tests and a credential-free smoke scenario.
</details>

<h3>Confidence Score: 5/5</h3>

This looks safe to merge.

No blocking issues found in the changed code. The provider lease path
releases resources on normal completion, stream close, and cancellation.
The admin apply path keeps restart-required changes separate from
provider-only hot swaps, and repository import paths appear updated for
the moved modules.

No files need follow-up attention.

<details><summary><h3><a href="https://www.greptile.com/trex"><img
alt="T-Rex"
src="https://greptile-static-assets.s3.amazonaws.com/trex/trex_green.svg"
height="20" align="absmiddle"></a> T-Rex Logs</h3></summary>

**What T-Rex did**
- Ran the T-Rex smoke test command and confirmed it completed with exit
code 0 and pytest passing.
- Monitored runtime ownership activity during the run, including a
provider A stream request on model-a generation 1, an admin publish to
generation 2, a new request on model-b generation 2, and the completion
of the original generation 1 stream.
- Collected smoke result artifacts from the .smoke-results area for
gateway 1, gateway 0, and main, and made them available for review.
- Opened the smoke report JSON artifacts to review the summarized
outcomes for each target environment.

<a
href="https://app.greptile.com/trex/runs/14008910/artifacts"><picture><source
media="(prefers-color-scheme: dark)"
srcset="https://greptile-static-assets.s3.amazonaws.com/badges/ViewAllArtifactsDark.svg?v=4"><source
media="(prefers-color-scheme: light)"
srcset="https://greptile-static-assets.s3.amazonaws.com/badges/ViewAllArtifacts.svg?v=4"><img
alt="View all artifacts"
src="https://greptile-static-assets.s3.amazonaws.com/badges/ViewAllArtifacts.svg?v=4"></picture></a>

<sub><a href="https://www.greptile.com/trex"><img alt="T-Rex"
src="https://greptile-static-assets.s3.amazonaws.com/trex/trex_green.svg"
height="14" align="absmiddle"></a> Ran code and verified through
T-Rex</sub>
</details>

<details open><summary><h3>Important Files Changed</h3></summary>

| Filename | Overview |
|----------|----------|
| src/free_claude_code/runtime/provider_manager.py | Adds provider
generation ownership, request leases, replacement, discovery refresh,
and shutdown cleanup. |
| src/free_claude_code/runtime/application.py | Adds the process-level
owner for startup, shutdown, admin operations, messaging, and session
control. |
| src/free_claude_code/api/routes.py | Routes now acquire provider
generation leases and bind them to response lifetime. |
| src/free_claude_code/api/response_streams.py | Adds response lifetime
binding so retained resources release after stream completion,
cancellation, or close. |
| src/free_claude_code/config/admin/persistence.py | Moves admin config
persistence into `config` and adds prepared validation plus atomic
managed-env commits. |
| src/free_claude_code/runtime/bootstrap.py | Adds the production
composition root for logging, runtime owners, services, and ASGI wiring.
|
| src/free_claude_code/api/__init__.py | Removes package-level API
re-exports as part of the HTTP adapter boundary cleanup. |

</details>

<details open><summary><h3>Sequence Diagram</h3></summary>

<a href="#gh-light-mode-only">

```mermaid
%%{init: {'theme': 'neutral'}}%%
sequenceDiagram
participant Client
participant API as FastAPI Route
participant Manager as ProviderRuntimeManager
participant Lease as Generation Lease
participant Runtime as Provider Generation
participant Admin as Admin Apply

Client->>API: Request /v1/messages or /v1/responses
API->>Manager: acquire()
Manager-->>API: lease for current generation
API->>Lease: resolve_provider()
Lease->>Runtime: use provider instance
Runtime-->>Client: response body or stream

Admin->>Manager: replace(candidate settings)
Manager->>Manager: publish new generation
Manager->>Manager: retire old generation

Client-->>API: response completes or disconnects
API->>Lease: release()
Lease->>Manager: decrement active leases
Manager->>Runtime: cleanup retired generation when drained
```

</a>
<a href="#gh-dark-mode-only">

```mermaid
%%{init: {'theme': 'base', 'themeVariables': {"darkMode": true, "background": "#0d1117", "primaryColor": "#21262d", "primaryTextColor": "#e6edf3", "primaryBorderColor": "#8b949e", "lineColor": "#8b949e", "textColor": "#e6edf3", "edgeLabelBackground": "#161b22", "actorBkg": "#21262d", "actorBorder": "#8b949e", "actorTextColor": "#e6edf3", "actorLineColor": "#8b949e", "signalColor": "#8b949e", "signalTextColor": "#e6edf3", "noteBkgColor": "#373320", "noteBorderColor": "#d4a72c", "noteTextColor": "#f0e6c0", "labelBoxBkgColor": "#21262d", "labelBoxBorderColor": "#8b949e", "labelTextColor": "#e6edf3", "loopTextColor": "#e6edf3", "activationBkgColor": "#30363d", "activationBorderColor": "#8b949e"}}}%%
sequenceDiagram
participant Client
participant API as FastAPI Route
participant Manager as ProviderRuntimeManager
participant Lease as Generation Lease
participant Runtime as Provider Generation
participant Admin as Admin Apply

Client->>API: Request /v1/messages or /v1/responses
API->>Manager: acquire()
Manager-->>API: lease for current generation
API->>Lease: resolve_provider()
Lease->>Runtime: use provider instance
Runtime-->>Client: response body or stream

Admin->>Manager: replace(candidate settings)
Manager->>Manager: publish new generation
Manager->>Manager: retire old generation

Client-->>API: response completes or disconnects
API->>Lease: release()
Lease->>Manager: decrement active leases
Manager->>Runtime: cleanup retired generation when drained
```

</a>
</details>

<sub>Reviews (1): Last reviewed commit: ["refactor: establish
single-owner
applica..."](https://github.com/alishahryar1/free-claude-code/commit/92fc06aa733b7acc34ad6ea50de8b6b4ce5cfac1)
| [Re-trigger
Greptile](https://app.greptile.com/api/retrigger?id=43349002)</sub>

<!-- /greptile_comment -->
2026-07-10 10:52:54 -07:00
Ali Khokhar 1278d00873 Restore protocol-correct provider errors without client retry loops (#1033)
## Problem

Provider failures before streaming were converted to HTTP 200 SSE
errors, masking typed upstream statuses as malformed proxy responses.
This regressed #1026's error visibility while solving client retry
loops; fixes #1031.

## Changes

| Before | After |
| --- | --- |
| Pre-start Messages failures returned HTTP 200 SSE errors. | Pre-start
Messages failures return typed non-2xx JSON with `x-should-retry:
false`. |
| Non-streaming Messages could preserve partial content after an
internal stream error. | Non-streaming Messages discard partial content
and return the mapped Anthropic error. |
| Post-start failures could synthesize a successful stop lifecycle. |
Post-start failures emit protocol-native terminal errors without a fake
success stop. |
| Responses failures lacked consistent retry ownership and correlation.
| Responses failures retain typed envelopes, retry suppression, response
IDs, and request IDs. |
| Request IDs were generated independently across layers. | One
ingress-owned request ID flows through response headers, provider calls,
error bodies, and traces. |
| Provider exceptions owned Anthropic serialization. | Neutral Anthropic
utilities own error envelopes, status mapping, and redacted diagnostics.
|
| Failure-path smoke expected the regressed HTTP 200 shape. |
Failure-path smoke verifies typed JSON and one downstream Claude CLI
request. |
| Package metadata remained at 3.4.15. | Package metadata and the
lockfile advance to 3.4.16. |

<!-- greptile_comment -->

<details open><summary><h3>Greptile Summary</h3></summary>

This PR restores protocol-correct provider error handling for Messages
and Responses. The main changes are:

- Pre-start provider failures return typed JSON errors with retry
suppression.
- Streaming failures use protocol-native terminal events after commit.
- Request IDs now flow from ingress through headers, traces, provider
calls, and error bodies.
- Shared Anthropic and OpenAI error helpers now shape payloads and
redact diagnostics.
- Provider-error smoke coverage and package metadata were updated.
</details>

<h3>Confidence Score: 4/5</h3>

Committed streaming failure paths still expose exception-derived
messages to clients.

Pre-start provider error handling is more protocol-correct, and the
completed-message lifecycle conflict appears fixed.

Messages and Responses streams still need fixed safe messages after the
HTTP response has committed.

<details><summary><h3><a href="https://www.greptile.com/trex"><img
alt="T-Rex"
src="https://greptile-static-assets.s3.amazonaws.com/trex/trex_green.svg"
height="20" align="absmiddle"></a> T-Rex Logs</h3></summary>

**What T-Rex did**
- I executed a focused uv-run harness against the Anthropic SSE
committed-stream path and reproduced a committed StreamingResponse that
yielded a terminal error SSE frame containing an internal diagnostic
marker.
- I ran a focused uv-run against iter\_responses\_sse\_from\_anthropic
and confirmed the stream was committed, with the later response.failed
SSE event carrying the internal diagnostic string in response.error.
- I completed provider error-handling verification, validating the
deterministic pytest path and running the local /v1/responses probe
script, which reported a 429 status with downstream\_call\_count: 1 and
related metadata, using the exact probe to exercise the failing
provider.

<a
href="https://app.greptile.com/trex/runs/13955393/artifacts"><picture><source
media="(prefers-color-scheme: dark)"
srcset="https://greptile-static-assets.s3.amazonaws.com/badges/ViewAllArtifactsDark.svg?v=4"><source
media="(prefers-color-scheme: light)"
srcset="https://greptile-static-assets.s3.amazonaws.com/badges/ViewAllArtifacts.svg?v=4"><img
alt="View all artifacts"
src="https://greptile-static-assets.s3.amazonaws.com/badges/ViewAllArtifacts.svg?v=4"></picture></a>

<sub><a href="https://www.greptile.com/trex"><img alt="T-Rex"
src="https://greptile-static-assets.s3.amazonaws.com/trex/trex_green.svg"
height="14" align="absmiddle"></a> Ran code and verified through
T-Rex</sub>
</details>

<details open><summary><h3>Important Files Changed</h3></summary>

| Filename | Overview |
|----------|----------|
| src/free_claude_code/api/response_streams.py | Pre-start Messages
errors now return typed JSON, but committed terminal SSE errors still
use exception-derived text. |
| src/free_claude_code/core/openai_responses/stream.py | Responses
streams now emit protocol failure events, but committed failures still
use exception-derived text. |
| src/free_claude_code/core/anthropic/streaming/ledger.py | Completed
Anthropic message streams now suppress late terminal errors after
`message_stop`. |
| src/free_claude_code/core/anthropic/errors.py | Shared Anthropic error
payload helpers now add request IDs, status mapping, and diagnostic
redaction. |

</details>

<a
href="https://app.greptile.com/api/ide/codex?prompt=IMPORTANT%3A%20Work%20in%20the%20repository%20%22alishahryar1%2Ffree-claude-code%22%20on%20the%20existing%20branch%20%22fix%2Fprotocol-correct-provider-errors%22.%20Checkout%20that%20branch%20%E2%80%94%20do%20NOT%20create%20a%20new%20branch%20or%20open%20a%20new%20PR.%20Push%20your%20changes%20to%20%22fix%2Fprotocol-correct-provider-errors%22.%0A%0AFix%20the%20following%202%20code%20review%20issues.%20Work%20through%20them%20one%20at%20a%20time%2C%20proposing%20concise%20fixes.%0A%0A---%0A%0A%23%23%23%20Issue%201%20of%202%0Asrc%2Ffree_claude_code%2Fapi%2Fresponse_streams.py%3A115-117%0A**Committed%20Stream%20Exception%20Text**%0A%0AWhen%20an%20Anthropic%20stream%20fails%20after%20the%20first%20chunk%20has%20committed%2C%20this%20terminal%20frame%20still%20derives%20the%20public%20SSE%20message%20from%20the%20exception.%20For%20unknown%20SDK%20or%20runtime%20errors%2C%20%60get_user_facing_error_message%28%29%60%20can%20return%20sanitized%20%60str%28exc%29%60%2C%20so%20internal%20URLs%2C%20payload%20fragments%2C%20or%20unsupported%20secret%20formats%20can%20still%20reach%20the%20client.%0A%0A%23%23%23%20Issue%202%20of%202%0Asrc%2Ffree_claude_code%2Fcore%2Fopenai_responses%2Fstream.py%3A49%0A**Failed%20Response%20Exception%20Text**%0A%0AAfter%20a%20Responses%20stream%20has%20emitted%20%60response.created%60%2C%20this%20failure%20event%20still%20uses%20exception-derived%20text%20as%20the%20public%20error%20message.%20If%20an%20unexpected%20provider%20or%20SDK%20exception%20includes%20internal%20diagnostics%20or%20a%20credential%20shape%20outside%20the%20redaction%20patterns%2C%20the%20committed%20%60response.failed%60%20event%20exposes%20it%20to%20the%20client.%0A%0A&repo=alishahryar1%2Ffree-claude-code&pr=1033&platform=github"><picture><source
media="(prefers-color-scheme: dark)"
srcset="https://greptile-static-assets.s3.amazonaws.com/badges/FixAllInCodexDark.svg?v=6"><source
media="(prefers-color-scheme: light)"
srcset="https://greptile-static-assets.s3.amazonaws.com/badges/FixAllInCodex.svg?v=6"><img
alt="Fix All in Codex"
src="https://greptile-static-assets.s3.amazonaws.com/badges/FixAllInCodex.svg?v=6"></picture></a>

<sub>Reviews (2): Last reviewed commit: ["fix(streaming): ignore errors
after
mess..."](https://github.com/alishahryar1/free-claude-code/commit/157eb504e9c8e15f6591bd72391810610be078f7)
| [Re-trigger
Greptile](https://app.greptile.com/api/retrigger?id=43242122)</sub>

> Greptile also left **2 inline comments** on this PR.

**Context used:**

- Context used - CLAUDE.md
([source](https://app.greptile.com/alishahryar1/github/Alishahryar1/free-claude-code/-/custom-context?memory=d2fd24d8-0dec-4faf-8ee4-e085e215a2f8))

<!-- /greptile_comment -->
2026-07-10 01:08:21 -07:00
Ali Khokhar 71a78a0c5a Move runtime packages under src namespace (#1029)
## Problem

Runtime modules were published as generic top-level packages like `api`,
`cli`, and `providers`. That shape is fragile for PyPI packaging and
weakens explicit ownership boundaries.

## Changes

| Before | After |
| --- | --- |
| Runtime code lived in root-level packages. | Runtime code lives under
`src/free_claude_code/`. |
| Console scripts targeted top-level modules. | Console scripts target
namespaced modules. |
| Tests and smoke helpers imported old package roots. | Tests and smoke
helpers import `free_claude_code.*`. |
| Packaging listed six root packages. | Packaging builds the single
namespaced package. |
| Contracts allowed old root package directories. | Contracts require
the src namespace and reject old root imports. |

<!-- greptile_comment -->

<details open><summary><h3>Greptile Summary</h3></summary>

This PR moves the runtime packages into the `src/free_claude_code`
namespace. The main changes are:

- Console scripts now point to `free_claude_code.*` entrypoints.
- Runtime imports, tests, and smoke helpers now use the namespaced
package.
- Packaging now builds the single `src/free_claude_code` package.
- Contract tests now reject old top-level runtime package roots and
imports.
</details>

<h3>Confidence Score: 5/5</h3>

This PR is safe to merge with minimal risk.

The changes are a broad but mostly mechanical namespace and
package-layout migration with updated packaging, tests, and contract
coverage.

No files require special attention.

<details><summary><h3><a href="https://www.greptile.com/trex"><img
alt="T-Rex"
src="https://greptile-static-assets.s3.amazonaws.com/trex/trex_green.svg"
height="20" align="absmiddle"></a> T-Rex Logs</h3></summary>

**What T-Rex did**
- Reviewed the primary contract validation by examining the namespace
validation log, which documents the exact commands executed, the working
directory, exit codes, pytest output, wheel build output, install
output, and import/entrypoint resolution.
- Verified the wheel listing by inspecting the wheel listing artifact,
confirming the available wheel filenames for the namespace validation.
- Ran and inspected the isolated import/entrypoint validation harness
saved as package-installed-import-check.py to validate import resolution
and entrypoint exposure.
- Captured and noted the wheel filename record in
package-wheel-filename.txt to enable traceability of the observed
artifact.

<a
href="https://app.greptile.com/trex/runs/13810533/artifacts"><picture><source
media="(prefers-color-scheme: dark)"
srcset="https://greptile-static-assets.s3.amazonaws.com/badges/ViewAllArtifactsDark.svg?v=4"><source
media="(prefers-color-scheme: light)"
srcset="https://greptile-static-assets.s3.amazonaws.com/badges/ViewAllArtifacts.svg?v=4"><img
alt="View all artifacts"
src="https://greptile-static-assets.s3.amazonaws.com/badges/ViewAllArtifacts.svg?v=4"></picture></a>

<sub><a href="https://www.greptile.com/trex"><img alt="T-Rex"
src="https://greptile-static-assets.s3.amazonaws.com/trex/trex_green.svg"
height="14" align="absmiddle"></a> Ran code and verified through
T-Rex</sub>
</details>

<details open><summary><h3>Important Files Changed</h3></summary>

| Filename | Overview |
|----------|----------|
| pyproject.toml | Updates packaging to build the single
`src/free_claude_code` package and retargets console scripts to
namespaced modules. |
| src/free_claude_code/config/env_template.py | Loads `.env.example`
from packaged resources with a source-checkout fallback after the
runtime package move. |
| src/free_claude_code/cli/entrypoints.py | Updates CLI entrypoint
imports to `free_claude_code.*` and continues to use the shared env
template loader. |
| src/free_claude_code/api/routes.py | Retargets API route dependencies
and handlers to the namespaced package without changing route behavior.
|
| src/free_claude_code/api/app.py | Updates app factory imports to the
namespaced package while preserving middleware, routers, and exception
handling. |
| src/free_claude_code/providers/runtime/factory.py | Updates lazy
provider factory imports to `free_claude_code.providers.*` under the new
package layout. |
| tests/contracts/test_import_boundaries.py | Adds contract coverage
requiring runtime packages to live under `src/free_claude_code` and
rejecting old top-level imports. |
| smoke/lib/child_process.py | Updates smoke child-process helpers to
import CLI entrypoints from the namespaced package. |
| README.md | Updates the project layout and extension guidance to refer
to `src/free_claude_code` and importable `free_claude_code.*` modules. |
| uv.lock | Reflects the package version bump associated with the
runtime packaging move. |

</details>

<details open><summary><h3>Sequence Diagram</h3></summary>

<a href="#gh-light-mode-only">

```mermaid
%%{init: {'theme': 'neutral'}}%%
sequenceDiagram
participant User as User / CLI
participant Script as Console script
participant Pkg as free_claude_code package
participant API as free_claude_code.api
participant Runtime as free_claude_code.providers.runtime
participant Provider as Provider adapter

User->>Script: run fcc-server / free-claude-code
Script->>Pkg: load free_claude_code.cli.entrypoints:serve
Pkg->>API: create FastAPI app and routes
API->>Runtime: resolve configured provider
Runtime->>Provider: instantiate namespaced adapter
Provider-->>Runtime: stream/model responses
Runtime-->>API: provider result
API-->>User: Anthropic/OpenAI-compatible response
```

</a>
<a href="#gh-dark-mode-only">

```mermaid
%%{init: {'theme': 'base', 'themeVariables': {"darkMode": true, "background": "#0d1117", "primaryColor": "#21262d", "primaryTextColor": "#e6edf3", "primaryBorderColor": "#8b949e", "lineColor": "#8b949e", "textColor": "#e6edf3", "edgeLabelBackground": "#161b22", "actorBkg": "#21262d", "actorBorder": "#8b949e", "actorTextColor": "#e6edf3", "actorLineColor": "#8b949e", "signalColor": "#8b949e", "signalTextColor": "#e6edf3", "noteBkgColor": "#373320", "noteBorderColor": "#d4a72c", "noteTextColor": "#f0e6c0", "labelBoxBkgColor": "#21262d", "labelBoxBorderColor": "#8b949e", "labelTextColor": "#e6edf3", "loopTextColor": "#e6edf3", "activationBkgColor": "#30363d", "activationBorderColor": "#8b949e"}}}%%
sequenceDiagram
participant User as User / CLI
participant Script as Console script
participant Pkg as free_claude_code package
participant API as free_claude_code.api
participant Runtime as free_claude_code.providers.runtime
participant Provider as Provider adapter

User->>Script: run fcc-server / free-claude-code
Script->>Pkg: load free_claude_code.cli.entrypoints:serve
Pkg->>API: create FastAPI app and routes
API->>Runtime: resolve configured provider
Runtime->>Provider: instantiate namespaced adapter
Provider-->>Runtime: stream/model responses
Runtime-->>API: provider result
API-->>User: Anthropic/OpenAI-compatible response
```

</a>
</details>

<sub>Reviews (2): Last reviewed commit: ["Fix documented package import
paths"](https://github.com/alishahryar1/free-claude-code/commit/bfa9f2704c45f3684da39657d5e13f3814e5d450)
| [Re-trigger
Greptile](https://app.greptile.com/api/retrigger?id=42950471)</sub>

<!-- /greptile_comment -->
2026-07-09 01:19:05 -07:00
Ali Khokhar a5310f4a0f Fix pre-start stream failures returning HTTP 200 (#1026)
## Problem

Provider streams could commit HTTP 200 before an upstream-backed first
SSE frame was available. When setup or retry failed before usable stream
output, Claude and Codex saw a successful but broken stream instead of a
retryable non-200 error.

## Changes

| Before | After |
| --- | --- |
| API egress returned `StreamingResponse` before probing the provider
iterator. | API egress waits for the first chunk before committing
success headers. |
| Pre-start provider failures became synthetic SSE success streams. |
Pre-start provider failures raise typed errors and return Anthropic or
OpenAI JSON with non-200 status. |
| Post-start unexpected stream failures could truncate protocol output.
| Post-start failures emit terminal Anthropic error or Responses
`response.failed` frames where possible. |
| Provider tests expected pre-start final failures as SSE tails. |
Provider tests assert typed pre-start errors and preserve midstream and
tool-salvage behavior. |

<!-- greptile_comment -->

<details open><summary><h3>Greptile Summary</h3></summary>

This PR changes streaming responses so HTTP success is not committed
before the first protocol chunk. The main changes are:

- First-chunk gated streaming helpers for Anthropic and OpenAI Responses
egress.
- Non-200 JSON error responses for provider failures before stream
output starts.
- Terminal Anthropic `error` and Responses `response.failed` frames for
post-start interruptions.
- Provider transport updates that raise typed pre-start errors while
preserving retry, recovery, and tool salvage paths.
- Targeted API and provider tests plus a patch version and lockfile
update.
</details>

<h3>Confidence Score: 5/5</h3>

Safe to merge with low risk.

The changed paths keep provider retry and recovery ownership in
transports, gate HTTP success before the first chunk, and preserve
cancellation behavior. Tests cover the main Anthropic and OpenAI
Responses pre-start and post-start failure paths.

No files require special attention.

<details><summary><h3><a href="https://www.greptile.com/trex"><img
alt="T-Rex"
src="https://greptile-static-assets.s3.amazonaws.com/trex/trex_green.svg"
height="20" align="absmiddle"></a> T-Rex Logs</h3></summary>

**What T-Rex did**
- Ran the pre-change stream-gating API tests against the previous commit
98515be55a, and they passed 35 tests in
5.08s (EXIT\_CODE: 0).
- Ran the post-change stream-gating API tests against the updated HEAD
d65f4ca107a024790df03cb4d30d6346027597b8, and they passed 35 tests in
2.27s (EXIT\_CODE: 0).
- Ran the focused API test suite against the current code and it passed
35 tests in 3.85s (EXIT\_CODE: 0).

<a
href="https://app.greptile.com/trex/runs/13779430/artifacts"><picture><source
media="(prefers-color-scheme: dark)"
srcset="https://greptile-static-assets.s3.amazonaws.com/badges/ViewAllArtifactsDark.svg?v=4"><source
media="(prefers-color-scheme: light)"
srcset="https://greptile-static-assets.s3.amazonaws.com/badges/ViewAllArtifacts.svg?v=4"><img
alt="View all artifacts"
src="https://greptile-static-assets.s3.amazonaws.com/badges/ViewAllArtifacts.svg?v=4"></picture></a>

<sub><a href="https://www.greptile.com/trex"><img alt="T-Rex"
src="https://greptile-static-assets.s3.amazonaws.com/trex/trex_green.svg"
height="14" align="absmiddle"></a> Ran code and verified through
T-Rex</sub>
</details>

<details open><summary><h3>Important Files Changed</h3></summary>

| Filename | Overview |
|----------|----------|
| api/response_streams.py | Introduces protocol-agnostic first-chunk
gating and Anthropic post-start terminal error fallback. |
| api/handlers/messages.py | Awaits first-chunk gated Anthropic stream
responses and maps pre-start provider or unexpected failures to JSON
errors. |
| api/handlers/responses.py | Adds first-chunk gated Responses streaming
and OpenAI-shaped pre-start error serialization. |
| providers/error_mapping.py | Adds shared mapping from final pre-start
stream exceptions to HTTP-serializable provider errors. |
| providers/transports/openai_chat/stream.py | Raises mapped provider
errors for uncommitted OpenAI-chat failures while preserving recovery
and complete tool salvage paths. |
| providers/transports/anthropic_messages/stream.py | Raises mapped
provider errors for uncommitted native stream failures while preserving
committed and salvageable tails. |
| core/openai_responses/stream.py | Preserves pre-start exceptions while
converting post-start Anthropic stream failures into same-assembler
`response.failed` events. |
| tests/api/test_response_streams.py | Adds unit coverage for
first-chunk gating, pre-start error JSON, and post-start terminal
Anthropic frames. |
| tests/api/test_openai_responses.py | Covers Responses pre-start
provider errors and same-id post-start failure frames. |
| tests/providers/test_openai_compat_5xx_retry.py | Covers exhausted
OpenAI-compatible 5xx and connection retries raising provider errors
before stream commit. |
| pyproject.toml | Bumps the package patch version for production
behavior changes. |
| uv.lock | Refreshes the lockfile package version to match the patch
bump. |

</details>

<details open><summary><h3>Sequence Diagram</h3></summary>

<a href="#gh-light-mode-only">

```mermaid
%%{init: {'theme': 'neutral'}}%%
sequenceDiagram
participant Client
participant Handler as API Handler
participant Egress as api/response_streams.py
participant Provider as Provider Stream
participant Assembler as Responses Assembler

Client->>Handler: POST /v1/messages or /v1/responses
Handler->>Provider: create async SSE iterator
Handler->>Egress: await first-chunk gated response
Egress->>Provider: anext(body)
alt Provider fails before first chunk
    Provider-->>Egress: ProviderError / exception
    Egress-->>Handler: protocol JSON error response
    Handler-->>Client: non-200 JSON
else First protocol chunk is available
    Provider-->>Egress: first SSE chunk
    Egress-->>Client: HTTP 200 StreamingResponse
    Egress-->>Client: replay first chunk and tail
    alt Anthropic post-start failure
        Provider-->>Egress: exception
        Egress-->>Client: terminal event: error
    else Responses post-start failure
        Provider-->>Assembler: exception after response.created
        Assembler-->>Client: response.failed with same response.id
    end
end
```

</a>
<a href="#gh-dark-mode-only">

```mermaid
%%{init: {'theme': 'base', 'themeVariables': {"darkMode": true, "background": "#0d1117", "primaryColor": "#21262d", "primaryTextColor": "#e6edf3", "primaryBorderColor": "#8b949e", "lineColor": "#8b949e", "textColor": "#e6edf3", "edgeLabelBackground": "#161b22", "actorBkg": "#21262d", "actorBorder": "#8b949e", "actorTextColor": "#e6edf3", "actorLineColor": "#8b949e", "signalColor": "#8b949e", "signalTextColor": "#e6edf3", "noteBkgColor": "#373320", "noteBorderColor": "#d4a72c", "noteTextColor": "#f0e6c0", "labelBoxBkgColor": "#21262d", "labelBoxBorderColor": "#8b949e", "labelTextColor": "#e6edf3", "loopTextColor": "#e6edf3", "activationBkgColor": "#30363d", "activationBorderColor": "#8b949e"}}}%%
sequenceDiagram
participant Client
participant Handler as API Handler
participant Egress as api/response_streams.py
participant Provider as Provider Stream
participant Assembler as Responses Assembler

Client->>Handler: POST /v1/messages or /v1/responses
Handler->>Provider: create async SSE iterator
Handler->>Egress: await first-chunk gated response
Egress->>Provider: anext(body)
alt Provider fails before first chunk
    Provider-->>Egress: ProviderError / exception
    Egress-->>Handler: protocol JSON error response
    Handler-->>Client: non-200 JSON
else First protocol chunk is available
    Provider-->>Egress: first SSE chunk
    Egress-->>Client: HTTP 200 StreamingResponse
    Egress-->>Client: replay first chunk and tail
    alt Anthropic post-start failure
        Provider-->>Egress: exception
        Egress-->>Client: terminal event: error
    else Responses post-start failure
        Provider-->>Assembler: exception after response.created
        Assembler-->>Client: response.failed with same response.id
    end
end
```

</a>
</details>

<sub>Reviews (1): Last reviewed commit: ["Fix pre-start stream failure
status"](https://github.com/alishahryar1/free-claude-code/commit/d65f4ca107a024790df03cb4d30d6346027597b8)
| [Re-trigger
Greptile](https://app.greptile.com/api/retrigger?id=42891757)</sub>

<!-- /greptile_comment -->
2026-07-08 18:15:41 -07:00
Ali Khokhar 745c38cbbe Move cloud providers to OpenAI-chat transport
Move remote cloud providers onto the OpenAI-chat transport and keep native Anthropic transport local-provider only.
2026-07-07 23:02:31 -07:00
Ali Khokhar dac6d4e88d Request streamed usage for OpenAI-chat providers (#1013)
## Problem

OpenAI-compatible streaming providers could return accurate final usage,
but FCC only requested it for DeepSeek and kept provider prompt tokens
out of final Anthropic usage.

## Changes

| Before | After |
| --- | --- |
| DeepSeek alone requested streamed usage. | The OpenAI-chat transport
requests streamed usage for all OpenAI-compatible providers. |
| Final input usage stayed on the local estimate even when providers
returned `prompt_tokens`. | Final input usage uses provider
`prompt_tokens` when available and falls back to the estimate when
absent. |
| Providers that rejected `stream_options.include_usage` failed the
request. | Providers that reject optional usage metadata retry once
without it. |
| DeepSeek owned duplicated usage extraction logic. | OpenAI-chat usage
extraction is shared, while DeepSeek keeps only cache-token mapping. |

<!-- greptile_comment -->

<details open><summary><h3>Greptile Summary</h3></summary>

This PR moves streamed usage handling into the shared OpenAI-chat
transport. The main changes are:

- Requests `stream_options.include_usage` for OpenAI-compatible
streaming providers.
- Uses provider `prompt_tokens` and `completion_tokens` when streamed
usage is returned.
- Falls back to local token estimates when usage metadata is absent or
rejected.
- Retries once without `include_usage` when an upstream provider rejects
the optional field.
- Keeps DeepSeek-specific cache token mapping in the DeepSeek provider.
</details>

<h3>Confidence Score: 5/5</h3>

Safe to merge with low risk.

The changed logic is localized to the OpenAI-chat streaming path and
includes focused regression coverage for the new usage and retry
behavior.

No files require special attention.

<details><summary><h3><a href="https://www.greptile.com/trex"><img
alt="T-Rex"
src="https://greptile-static-assets.s3.amazonaws.com/trex/trex_green.svg"
height="20" align="absmiddle"></a> T-Rex Logs</h3></summary>

**What T-Rex did**
- The team executed the general contract validation for the non-UI
provider transport using a mocked chat-completion streaming surface.
- Runtime artifacts, including the runtime log and before/after
comparison artifacts, were produced for inspection.
- The after-run summary shows 10 passed in 5.38 seconds with EXIT\_CODE
0.
- Because the test used a mocked streaming surface rather than live
endpoints, no HTTP endpoints or HTTP status messages were observed.

<a
href="https://app.greptile.com/trex/runs/13637854/artifacts"><picture><source
media="(prefers-color-scheme: dark)"
srcset="https://greptile-static-assets.s3.amazonaws.com/badges/ViewAllArtifactsDark.svg?v=4"><source
media="(prefers-color-scheme: light)"
srcset="https://greptile-static-assets.s3.amazonaws.com/badges/ViewAllArtifacts.svg?v=4"><img
alt="View all artifacts"
src="https://greptile-static-assets.s3.amazonaws.com/badges/ViewAllArtifacts.svg?v=4"></picture></a>

<sub><a href="https://www.greptile.com/trex"><img alt="T-Rex"
src="https://greptile-static-assets.s3.amazonaws.com/trex/trex_green.svg"
height="14" align="absmiddle"></a> Ran code and verified through
T-Rex</sub>
</details>

<details open><summary><h3>Important Files Changed</h3></summary>

| Filename | Overview |
|----------|----------|
| core/anthropic/streaming/ledger.py | Allows final `message_delta`
usage to override the ledger's initial estimated input token count. |
| providers/deepseek/client.py | Reuses shared usage extraction for
DeepSeek cache hit and miss token mapping. |
| providers/transports/openai_chat/stream.py | Requests streamed usage
and uses provider prompt and completion tokens when present. |
| providers/transports/openai_chat/transport.py | Adds bounded retry
behavior when upstream providers reject streamed usage metadata. |
| providers/transports/openai_chat/usage.py | Adds shared helpers for
usage requests, usage extraction, and usage rejection detection. |
| tests/providers/test_openai_chat_usage.py | Adds coverage for streamed
usage helpers, provider token usage, and retry fallback behavior. |

</details>

<details open><summary><h3>Sequence Diagram</h3></summary>

<a href="#gh-light-mode-only">

```mermaid
%%{init: {'theme': 'neutral'}}%%
sequenceDiagram
participant Client as Anthropic client
participant Adapter as OpenAIChatStreamAdapter
participant Transport as OpenAIChatTransport
participant Provider as OpenAI-compatible provider
participant Ledger as AnthropicStreamLedger

Client->>Adapter: stream_response(request, input_tokens)
Adapter->>Adapter: build body + request_stream_usage()
Adapter->>Transport: _create_stream(body with include_usage)
Transport->>Provider: "chat.completions.create(stream=True)"
alt provider rejects include_usage
    Provider-->>Transport: 400/422 usage option error
    Transport->>Transport: clone_without_stream_usage()
    Transport->>Provider: retry without include_usage
end
Provider-->>Adapter: streaming chunks + optional final usage
Adapter->>Adapter: usage_int(prompt_tokens/completion_tokens)
Adapter->>Ledger: message_delta(input_tokens, output_tokens, provider usage fields)
Ledger-->>Client: Anthropic SSE usage
```

</a>
<a href="#gh-dark-mode-only">

```mermaid
%%{init: {'theme': 'base', 'themeVariables': {"darkMode": true, "background": "#0d1117", "primaryColor": "#21262d", "primaryTextColor": "#e6edf3", "primaryBorderColor": "#8b949e", "lineColor": "#8b949e", "textColor": "#e6edf3", "edgeLabelBackground": "#161b22", "actorBkg": "#21262d", "actorBorder": "#8b949e", "actorTextColor": "#e6edf3", "actorLineColor": "#8b949e", "signalColor": "#8b949e", "signalTextColor": "#e6edf3", "noteBkgColor": "#373320", "noteBorderColor": "#d4a72c", "noteTextColor": "#f0e6c0", "labelBoxBkgColor": "#21262d", "labelBoxBorderColor": "#8b949e", "labelTextColor": "#e6edf3", "loopTextColor": "#e6edf3", "activationBkgColor": "#30363d", "activationBorderColor": "#8b949e"}}}%%
sequenceDiagram
participant Client as Anthropic client
participant Adapter as OpenAIChatStreamAdapter
participant Transport as OpenAIChatTransport
participant Provider as OpenAI-compatible provider
participant Ledger as AnthropicStreamLedger

Client->>Adapter: stream_response(request, input_tokens)
Adapter->>Adapter: build body + request_stream_usage()
Adapter->>Transport: _create_stream(body with include_usage)
Transport->>Provider: "chat.completions.create(stream=True)"
alt provider rejects include_usage
    Provider-->>Transport: 400/422 usage option error
    Transport->>Transport: clone_without_stream_usage()
    Transport->>Provider: retry without include_usage
end
Provider-->>Adapter: streaming chunks + optional final usage
Adapter->>Adapter: usage_int(prompt_tokens/completion_tokens)
Adapter->>Ledger: message_delta(input_tokens, output_tokens, provider usage fields)
Ledger-->>Client: Anthropic SSE usage
```

</a>
</details>

<sub>Reviews (1): Last reviewed commit: ["Request streamed usage for
OpenAI-chat
p..."](https://github.com/alishahryar1/free-claude-code/commit/0473d8fe4bd042b5e247fc1dfa479afed9b60529)
| [Re-trigger
Greptile](https://app.greptile.com/api/retrigger?id=42593974)</sub>

<!-- /greptile_comment -->
2026-07-07 19:51:55 -07:00
Ali Khokhar ccf46b88cf Retry pre-stream provider transport failures (#1003) 2026-07-07 00:13:47 -07:00
Ali Khokhar bd85deb736 Fix OpenAI chat reasoning and tool history replay (#1002)
## Problem

OpenAI-chat providers lost explicit empty reasoning state and could
replay invalid tool-call history when unrelated messages appeared before
matching tool results.

## Changes

| Before | After |
| --- | --- |
| Empty `reasoning_content` and empty thinking blocks were treated as
absent. | Empty reasoning is preserved as explicit replay state. |
| OpenAI-chat conversion only deferred post-tool assistant text. |
OpenAI-chat conversion buffers later transcript messages until required
tool results are emitted. |
| Responses prior tool calls and outputs were emitted one item per
message. | Responses prior tool calls and outputs are grouped into valid
Anthropic tool-use/result messages. |
| Empty streamed `reasoning_content` produced no thinking block. | Empty
streamed `reasoning_content` starts thinking state without visible delta
text. |

<!-- greptile_comment -->

<details open><summary><h3>Greptile Summary</h3></summary>

This PR fixes OpenAI chat reasoning replay and tool-history ordering.
The main changes are:

- Preserves explicit empty `reasoning_content` and empty thinking
blocks.
- Reworks OpenAI chat conversion around a ledger that waits for required
tool results before replaying buffered transcript messages.
- Groups prior Responses tool calls and outputs into valid Anthropic
tool-use and tool-result messages.
- Starts streamed thinking state when empty `reasoning_content` is
received.
- Adds focused tests for nested tool turns, out-of-order results,
multi-tool replay, and empty reasoning.
</details>

<h3>Confidence Score: 5/5</h3>

Safe to merge with low risk.

No blocking issues were found in the changed conversion paths. The
updated ledger covers the prior invalid replay cases and the tests
include nested, out-of-order, multi-tool, and empty reasoning scenarios.
The required patch version and lockfile updates are present.

No files require special attention.

<details><summary><h3><a href="https://www.greptile.com/trex"><img
alt="T-Rex"
src="https://greptile-static-assets.s3.amazonaws.com/trex/trex_green.svg"
height="20" align="absmiddle"></a> T-Rex Logs</h3></summary>

**What T-Rex did**
- Ran the focused OpenAI conversion regression suite with Pytest,
capturing the command, working directory, pass count, exit code, and
elapsed time.
- Encountered an external timeout during the initial Pytest run at 98%
progress, then re-ran the same focused suite to completion for
definitive proof.
- Validated code quality with Ruff by executing the lint command and
obtaining a successful output.

<a
href="https://app.greptile.com/trex/runs/13503694/artifacts"><picture><source
media="(prefers-color-scheme: dark)"
srcset="https://greptile-static-assets.s3.amazonaws.com/badges/ViewAllArtifactsDark.svg?v=4"><source
media="(prefers-color-scheme: light)"
srcset="https://greptile-static-assets.s3.amazonaws.com/badges/ViewAllArtifacts.svg?v=4"><img
alt="View all artifacts"
src="https://greptile-static-assets.s3.amazonaws.com/badges/ViewAllArtifacts.svg?v=4"></picture></a>

<sub><a href="https://www.greptile.com/trex"><img alt="T-Rex"
src="https://greptile-static-assets.s3.amazonaws.com/trex/trex_green.svg"
height="14" align="absmiddle"></a> Ran code and verified through
T-Rex</sub>
</details>

<details open><summary><h3>Important Files Changed</h3></summary>

| Filename | Overview |
|----------|----------|
| core/anthropic/conversion.py | Replaces single pending-tool state with
a ledger that buffers transcript segments until required OpenAI chat
tool results can be emitted in valid order. |
| core/openai_responses/input.py | Groups consecutive prior Responses
tool calls/results into Anthropic tool-use/result turns and preserves
explicit empty reasoning. |
| core/openai_responses/reasoning.py | Updates reasoning extraction and
combination helpers so empty strings remain explicit replay state
without adding spurious separators. |
| providers/deepseek/compat.py | Treats empty top-level or block-level
thinking as replayable when detecting DeepSeek tool-history
compatibility. |
| providers/transports/openai_chat/stream.py | Starts an Anthropic
thinking block for empty streamed `reasoning_content` while only
emitting deltas for non-empty text. |
| tests/providers/test_converter.py | Adds OpenAI chat conversion
coverage for buffered tool history, nested pending tool turns, and
explicit empty reasoning. |
| tests/core/openai_responses/test_conversion.py | Adds Responses
conversion tests for grouped prior tool calls/results and empty
reasoning attachment. |
| pyproject.toml | Bumps the package patch version for the production
conversion fixes. |
| uv.lock | Keeps the lockfile package version in sync with
`pyproject.toml`. |

</details>

<details open><summary><h3>Sequence Diagram</h3></summary>

<a href="#gh-light-mode-only">

```mermaid
%%{init: {'theme': 'neutral'}}%%
sequenceDiagram
participant A as Anthropic transcript
participant L as OpenAI chat ledger
participant O as OpenAI chat history
A->>L: Assistant tool_use segment
L->>O: Emit assistant tool_calls
A->>L: Later plain user/assistant messages
L-->>L: Buffer until required tool_result ids arrive
A->>L: User tool_result blocks
L->>O: Emit matching role: tool results in tool_call order
L->>O: Emit deferred assistant post-tool content
L->>O: Drain buffered plain transcript messages
```

</a>
<a href="#gh-dark-mode-only">

```mermaid
%%{init: {'theme': 'base', 'themeVariables': {"darkMode": true, "background": "#0d1117", "primaryColor": "#21262d", "primaryTextColor": "#e6edf3", "primaryBorderColor": "#8b949e", "lineColor": "#8b949e", "textColor": "#e6edf3", "edgeLabelBackground": "#161b22", "actorBkg": "#21262d", "actorBorder": "#8b949e", "actorTextColor": "#e6edf3", "actorLineColor": "#8b949e", "signalColor": "#8b949e", "signalTextColor": "#e6edf3", "noteBkgColor": "#373320", "noteBorderColor": "#d4a72c", "noteTextColor": "#f0e6c0", "labelBoxBkgColor": "#21262d", "labelBoxBorderColor": "#8b949e", "labelTextColor": "#e6edf3", "loopTextColor": "#e6edf3", "activationBkgColor": "#30363d", "activationBorderColor": "#8b949e"}}}%%
sequenceDiagram
participant A as Anthropic transcript
participant L as OpenAI chat ledger
participant O as OpenAI chat history
A->>L: Assistant tool_use segment
L->>O: Emit assistant tool_calls
A->>L: Later plain user/assistant messages
L-->>L: Buffer until required tool_result ids arrive
A->>L: User tool_result blocks
L->>O: Emit matching role: tool results in tool_call order
L->>O: Emit deferred assistant post-tool content
L->>O: Drain buffered plain transcript messages
```

</a>
</details>

<sub>Reviews (3): Last reviewed commit: ["Refactor OpenAI chat tool
history
replay"](https://github.com/alishahryar1/free-claude-code/commit/ae1635d2ce3a232ba7f4c0b9787f7b604b625544)
| [Re-trigger
Greptile](https://app.greptile.com/api/retrigger?id=42274270)</sub>

<!-- /greptile_comment -->
2026-07-06 23:00:06 -07:00
Ali Khokhar 950aba393d Disable Hugging Face reasoning replay
Disable Hugging Face prior reasoning replay for Chat Completions while preserving streamed reasoning output.
2026-07-06 20:25:07 -07:00
Ali Khokhar 62c0480eed Add Mistral reasoning fallback 2026-07-06 01:46:20 -07:00
Ali Khokhar 47ddedcada Fix NIM chat template downgrade (#997)
## Problem

NVIDIA NIM Mistral-tokenizer models can reject chat-template controls
with HTTP 400. FCC only removed `chat_template`, so requests that only
had `chat_template_kwargs` still failed.

## Changes

| Before | After |
| --- | --- |
| NIM retried chat-template errors by stripping only
`extra_body.chat_template`. | NIM retries chat-template errors by
stripping `extra_body.chat_template` and
`extra_body.chat_template_kwargs`. |
| Mistral-tokenizer models could fail after rejecting thinking
chat-template kwargs. | Mistral-tokenizer models retry once without NIM
chat-template controls. |
| Reasoning-budget downgrades preserved thinking flags independently. |
Reasoning-budget downgrades still preserve thinking flags independently.
|
| Package metadata stayed at `3.4.2`. | Package metadata is bumped to
`3.4.3`. |

<!-- greptile_comment -->

<details open><summary><h3>Greptile Summary</h3></summary>

This PR fixes the NVIDIA NIM chat-template retry downgrade. The main
changes are:

- Removes both `extra_body.chat_template` and
`extra_body.chat_template_kwargs` before retrying chat-template 400
errors.
- Adds tests for the full chat-template path and the kwargs-only
regression case.
- Adds helper coverage for unchanged request bodies.
- Bumps package metadata from `3.4.2` to `3.4.3` and keeps `uv.lock`
aligned.
</details>

<h3>Confidence Score: 5/5</h3>

Safe to merge with minimal risk.

The production change is narrow and covered by targeted tests for both
chat-template stripping paths. Version metadata and lockfile updates
follow the repository guidance.

No files require special attention.

<details><summary><h3><a href="https://www.greptile.com/trex"><img
alt="T-Rex"
src="https://greptile-static-assets.s3.amazonaws.com/trex/trex_green.svg"
height="20" align="absmiddle"></a> T-Rex Logs</h3></summary>

**What T-Rex did**
- The initial focused chat-template test run was executed and showed 6
tests passed, while the outer shell wrapper returned nonzero due to a
PIPESTATUS\[0\] substitution issue under /bin/sh after pytest completed.
- A corrected wrapper rerun of the same focused chat-template tests was
performed, and 6 tests passed with exit code 0.
- Focused provider retry tests for chat\_template or reasoning\_budget
were executed, and 4 tests passed with exit code 0.

<a
href="https://app.greptile.com/trex/runs/13364237/artifacts"><picture><source
media="(prefers-color-scheme: dark)"
srcset="https://greptile-static-assets.s3.amazonaws.com/badges/ViewAllArtifactsDark.svg?v=4"><source
media="(prefers-color-scheme: light)"
srcset="https://greptile-static-assets.s3.amazonaws.com/badges/ViewAllArtifacts.svg?v=4"><img
alt="View all artifacts"
src="https://greptile-static-assets.s3.amazonaws.com/badges/ViewAllArtifacts.svg?v=4"></picture></a>

<sub><a href="https://www.greptile.com/trex"><img alt="T-Rex"
src="https://greptile-static-assets.s3.amazonaws.com/trex/trex_green.svg"
height="14" align="absmiddle"></a> Ran code and verified through
T-Rex</sub>
</details>

<details open><summary><h3>Important Files Changed</h3></summary>

| Filename | Overview |
|----------|----------|
| providers/nvidia_nim/retry.py | Extends NIM chat-template retry
downgrades to remove both `chat_template` and `chat_template_kwargs`
while preserving unchanged-body behavior. |
| tests/providers/test_nvidia_nim.py | Updates streaming retry coverage
to assert the second request strips all chat-template controls,
including the kwargs-only regression case. |
| tests/providers/test_nvidia_nim_request.py | Adds request-body helper
coverage for stripping `chat_template_kwargs` and returning `None` when
no chat-template controls are present. |
| pyproject.toml | Bumps package version from `3.4.2` to `3.4.3` for the
production bug fix. |
| uv.lock | Keeps the editable package version in the lockfile aligned
with `pyproject.toml`. |

</details>

<details open><summary><h3>Sequence Diagram</h3></summary>

<a href="#gh-light-mode-only">

```mermaid
%%{init: {'theme': 'neutral'}}%%
sequenceDiagram
participant Client as Claude request
participant Provider as NvidiaNimProvider
participant NIM as NVIDIA NIM API

Client->>Provider: stream_response(request)
Provider->>NIM: create(extra_body with chat_template controls)
NIM-->>Provider: HTTP 400 chat_template rejection
Provider->>Provider: clone_body_without_chat_template()
Provider->>Provider: remove chat_template and chat_template_kwargs
Provider->>NIM: retry create(extra_body without chat-template controls)
NIM-->>Provider: stream chunks
Provider-->>Client: SSE response events
```

</a>
<a href="#gh-dark-mode-only">

```mermaid
%%{init: {'theme': 'base', 'themeVariables': {"darkMode": true, "background": "#0d1117", "primaryColor": "#21262d", "primaryTextColor": "#e6edf3", "primaryBorderColor": "#8b949e", "lineColor": "#8b949e", "textColor": "#e6edf3", "edgeLabelBackground": "#161b22", "actorBkg": "#21262d", "actorBorder": "#8b949e", "actorTextColor": "#e6edf3", "actorLineColor": "#8b949e", "signalColor": "#8b949e", "signalTextColor": "#e6edf3", "noteBkgColor": "#373320", "noteBorderColor": "#d4a72c", "noteTextColor": "#f0e6c0", "labelBoxBkgColor": "#21262d", "labelBoxBorderColor": "#8b949e", "labelTextColor": "#e6edf3", "loopTextColor": "#e6edf3", "activationBkgColor": "#30363d", "activationBorderColor": "#8b949e"}}}%%
sequenceDiagram
participant Client as Claude request
participant Provider as NvidiaNimProvider
participant NIM as NVIDIA NIM API

Client->>Provider: stream_response(request)
Provider->>NIM: create(extra_body with chat_template controls)
NIM-->>Provider: HTTP 400 chat_template rejection
Provider->>Provider: clone_body_without_chat_template()
Provider->>Provider: remove chat_template and chat_template_kwargs
Provider->>NIM: retry create(extra_body without chat-template controls)
NIM-->>Provider: stream chunks
Provider-->>Client: SSE response events
```

</a>
</details>

<sub>Reviews (1): Last reviewed commit: ["Fix NIM chat template
downgrade"](https://github.com/alishahryar1/free-claude-code/commit/fb2d152372356e22441bd52aca99b1dca4792b61)
| [Re-trigger
Greptile](https://app.greptile.com/api/retrigger?id=42001002)</sub>

<!-- /greptile_comment -->
2026-07-05 22:56:45 -07:00
Ali Khokhar 15cab79a43 Update GLM 5.2 model references 2026-07-05 18:28:58 -07:00
newmemories360 770d56708a Add SambaNova Cloud provider (#990) 2026-07-05 12:53:05 -07:00
newmemories360 418f4963e5 Fix HTTP 400 when max_(completion_)tokens exceeds a model's cap (#955) (#991) 2026-07-05 12:19:36 -07:00
Ali Khokhar 0b86dd4ef8 Add GitHub Models provider (#989)
## Problem

FCC does not expose GitHub Models, so users with GitHub Models access
cannot route Claude, Codex, or messaging prompts through GitHub's hosted
model catalog.

## Changes

| Before | After |
| --- | --- |
| Provider catalog did not include GitHub Models. | Provider catalog
includes `github_models` with token, proxy, admin, smoke, and model
picker wiring. |
| Requests could not target GitHub Models inference. |
`providers/github_models` routes OpenAI-chat requests to
`https://models.github.ai/inference`. |
| Model discovery assumed provider `/models` compatibility. | GitHub
Models discovery uses the catalog API and advertises stream/tool-capable
models. |
| OpenAI-chat transport could not set provider default headers. |
OpenAI-chat transport accepts provider-owned default headers. |
| Docs and templates omitted GitHub Models setup. | README,
`.env.example`, and architecture docs document GitHub Models setup and
ownership. |

<!-- greptile_comment -->

<details open><summary><h3>Greptile Summary</h3></summary>

This PR adds GitHub Models as a new provider. The main changes are:

- New `github_models` provider runtime, catalog, settings, and admin
wiring.
- OpenAI-chat transport support for provider-owned default headers.
- GitHub Models catalog discovery filtered to streaming and tool-capable
models.
- Smoke configuration, environment template, docs, and tests for the new
provider.
- Package version and lockfile updates for the new feature.
</details>

<h3>Confidence Score: 5/5</h3>

Safe to merge with low risk.

The provider is wired through runtime creation, catalog metadata,
settings, admin fields, smoke config, docs, version metadata, and
focused tests. No blocking correctness or security issues were found in
the changed paths.

No files require special attention.

<details><summary><h3><a href="https://www.greptile.com/trex"><img
alt="T-Rex"
src="https://greptile-static-assets.s3.amazonaws.com/trex/trex_green.svg"
height="20" align="absmiddle"></a> T-Rex Logs</h3></summary>

**What T-Rex did**
- Before-change focused pytest run against HEAD^ showed no GitHub Models
provider tests were collected.
- After-change focused pytest run showed all 37 provider/runtime tests
passed.
- After-change harness output captured structured evidence for catalog
discovery and OpenAI-chat request routing, and the harness exited
successfully.
- A temporary harness Python script was generated to capture the mocked
request/response evidence.

<a
href="https://app.greptile.com/trex/runs/13317138/artifacts"><picture><source
media="(prefers-color-scheme: dark)"
srcset="https://greptile-static-assets.s3.amazonaws.com/badges/ViewAllArtifactsDark.svg?v=4"><source
media="(prefers-color-scheme: light)"
srcset="https://greptile-static-assets.s3.amazonaws.com/badges/ViewAllArtifacts.svg?v=4"><img
alt="View all artifacts"
src="https://greptile-static-assets.s3.amazonaws.com/badges/ViewAllArtifacts.svg?v=4"></picture></a>

<sub><a href="https://www.greptile.com/trex"><img alt="T-Rex"
src="https://greptile-static-assets.s3.amazonaws.com/trex/trex_green.svg"
height="14" align="absmiddle"></a> Ran code and verified through
T-Rex</sub>
</details>

<details open><summary><h3>Important Files Changed</h3></summary>

| Filename | Overview |
|----------|----------|
| providers/github_models/client.py | Implements GitHub Models
OpenAI-chat transport wiring, default GitHub headers, and catalog-based
stream/tool-capable model discovery. |
| providers/transports/openai_chat/transport.py | Allows OpenAI-chat
providers to pass default headers into the shared AsyncOpenAI client. |
| config/provider_catalog.py | Registers GitHub Models provider
metadata, default inference base URL, credential, proxy, and
capabilities. |
| config/settings.py | Adds settings bindings for `GITHUB_MODELS_TOKEN`
and `GITHUB_MODELS_PROXY`. |
| api/admin_config/provider_manifest.py | Adds GitHub Models token
labeling and description for generated admin provider fields. |
| smoke/lib/config.py | Adds GitHub Models smoke defaults and credential
detection. |
| tests/providers/test_github_models.py | Adds focused tests for GitHub
Models initialization, request conversion, catalog filtering, streaming,
tool calls, reasoning, and cleanup. |
| tests/providers/test_provider_runtime.py | Covers GitHub Models
descriptor, provider config construction, and runtime instantiation. |
| README.md | Adds GitHub Models setup documentation and updates
provider counts/numbering. |
| pyproject.toml | Bumps the package version to `3.2.0` for the new
provider feature. |

</details>

<details open><summary><h3>Sequence Diagram</h3></summary>

<a href="#gh-light-mode-only">

```mermaid
%%{init: {'theme': 'neutral'}}%%
sequenceDiagram
participant User as Claude/Codex client
participant FCC as FCC proxy/router
participant Factory as Provider runtime factory
participant GH as GitHubModelsProvider
participant OpenAI as Shared OpenAI-chat transport
participant API as models.github.ai

User->>FCC: Request with model `github_models/...`
FCC->>Factory: create_provider(`github_models`, settings)
Factory->>GH: ProviderConfig(token, base_url, proxy)
GH->>OpenAI: Initialize with GitHub default headers
FCC->>GH: stream_response(MessagesRequest)
GH->>OpenAI: build OpenAI chat body
OpenAI->>API: "POST /inference/chat/completions (stream=true)"
API-->>OpenAI: OpenAI-compatible stream chunks
OpenAI-->>FCC: Anthropic SSE events
FCC-->>User: Streamed Anthropic response

FCC->>GH: list_model_infos()
GH->>API: GET /catalog/models
API-->>GH: Catalog entries with capabilities
GH-->>FCC: stream/tool-capable model ids
```

</a>
<a href="#gh-dark-mode-only">

```mermaid
%%{init: {'theme': 'base', 'themeVariables': {"darkMode": true, "background": "#0d1117", "primaryColor": "#21262d", "primaryTextColor": "#e6edf3", "primaryBorderColor": "#8b949e", "lineColor": "#8b949e", "textColor": "#e6edf3", "edgeLabelBackground": "#161b22", "actorBkg": "#21262d", "actorBorder": "#8b949e", "actorTextColor": "#e6edf3", "actorLineColor": "#8b949e", "signalColor": "#8b949e", "signalTextColor": "#e6edf3", "noteBkgColor": "#373320", "noteBorderColor": "#d4a72c", "noteTextColor": "#f0e6c0", "labelBoxBkgColor": "#21262d", "labelBoxBorderColor": "#8b949e", "labelTextColor": "#e6edf3", "loopTextColor": "#e6edf3", "activationBkgColor": "#30363d", "activationBorderColor": "#8b949e"}}}%%
sequenceDiagram
participant User as Claude/Codex client
participant FCC as FCC proxy/router
participant Factory as Provider runtime factory
participant GH as GitHubModelsProvider
participant OpenAI as Shared OpenAI-chat transport
participant API as models.github.ai

User->>FCC: Request with model `github_models/...`
FCC->>Factory: create_provider(`github_models`, settings)
Factory->>GH: ProviderConfig(token, base_url, proxy)
GH->>OpenAI: Initialize with GitHub default headers
FCC->>GH: stream_response(MessagesRequest)
GH->>OpenAI: build OpenAI chat body
OpenAI->>API: "POST /inference/chat/completions (stream=true)"
API-->>OpenAI: OpenAI-compatible stream chunks
OpenAI-->>FCC: Anthropic SSE events
FCC-->>User: Streamed Anthropic response

FCC->>GH: list_model_infos()
GH->>API: GET /catalog/models
API-->>GH: Catalog entries with capabilities
GH-->>FCC: stream/tool-capable model ids
```

</a>
</details>

<sub>Reviews (1): Last reviewed commit: ["Add GitHub Models
provider"](https://github.com/alishahryar1/free-claude-code/commit/736d3f9213f6a8d243d4001c5135b7fac402f143)
| [Re-trigger
Greptile](https://app.greptile.com/api/retrigger?id=41901295)</sub>

<!-- /greptile_comment -->
2026-07-05 12:11:55 -07:00
Ali Khokhar 9a17d1ed0a Add Cohere provider (#986)
## Problem

FCC does not expose Cohere's OpenAI-compatible chat models, so users
with Cohere keys cannot route Claude, Codex, or messaging prompts
through Cohere.

## Changes

| Before | After |
| --- | --- |
| Provider catalog did not include Cohere. | Provider catalog includes
Cohere with `COHERE_API_KEY`, `COHERE_PROXY`, admin status, and smoke
model wiring. |
| Requests could not target Cohere's compatibility API. |
`providers/cohere` routes OpenAI-chat requests to Cohere's compatibility
API with Cohere-specific request policy. |
| Docs and templates omitted Cohere setup. | README, `.env.example`, and
architecture docs document Cohere setup and ownership. |
| Cohere behavior had no regression coverage. | Provider, runtime,
admin, config, smoke, and catalog tests cover Cohere integration. |

<!-- greptile_comment -->

<details open><summary><h3>Greptile Summary</h3></summary>

This PR adds Cohere as a new OpenAI-compatible chat provider. The main
changes are:

- Cohere provider metadata in the catalog, settings, Admin UI manifest,
and runtime factory.
- A new `CohereProvider` using the shared OpenAI chat transport with
Cohere-specific request shaping.
- Cohere API key, proxy, smoke model, README, architecture, and
environment template updates.
- Tests for Admin config, settings, provider catalog order, smoke
config, runtime creation, and Cohere request/stream behavior.
- Version and lockfile updates for the new provider feature.
</details>

<h3>Confidence Score: 5/5</h3>

Safe to merge with minimal risk.

No functional, security, or contract issues were identified. Cohere is
consistently wired through settings, catalog metadata, factory creation,
Admin config, smoke defaults, docs, versioning, and targeted tests. The
implemented Cohere `reasoning_effort` values match the Compatibility API
behavior checked during review.

No files require special attention.

<details><summary><h3><a href="https://www.greptile.com/trex"><img
alt="T-Rex"
src="https://greptile-static-assets.s3.amazonaws.com/trex/trex_green.svg"
height="20" align="absmiddle"></a> T-Rex Logs</h3></summary>

**What T-Rex did**
- Validated the provider runtime handling of Cohere requests, including
the request body policy, streaming parsing, and default base URL and API
key behavior.
- Verified that the runtime descriptor wiring and provider config
proxy/key behavior pass in the general contract validation.
- Confirmed the admin/config smoke contract artifact shows the Cohere
environment settings, admin config masking, feature/provider catalog
contracts, and smoke configuration passing.
- Compared the initial -01-before.log and the clean -02-after.log
captures to confirm the same scoped commands are present and both exit
with code 0.

<a
href="https://app.greptile.com/trex/runs/13309850/artifacts"><picture><source
media="(prefers-color-scheme: dark)"
srcset="https://greptile-static-assets.s3.amazonaws.com/badges/ViewAllArtifactsDark.svg?v=4"><source
media="(prefers-color-scheme: light)"
srcset="https://greptile-static-assets.s3.amazonaws.com/badges/ViewAllArtifacts.svg?v=4"><img
alt="View all artifacts"
src="https://greptile-static-assets.s3.amazonaws.com/badges/ViewAllArtifacts.svg?v=4"></picture></a>

<sub><a href="https://www.greptile.com/trex"><img alt="T-Rex"
src="https://greptile-static-assets.s3.amazonaws.com/trex/trex_green.svg"
height="14" align="absmiddle"></a> Ran code and verified through
T-Rex</sub>
</details>

<details open><summary><h3>Important Files Changed</h3></summary>

| Filename | Overview |
|----------|----------|
| providers/cohere/client.py | Implements Cohere request shaping over
shared OpenAI chat transport, including allowed extra body and reasoning
mapping; no issues found. |
| config/provider_catalog.py | Registers Cohere with credential, proxy,
base URL, transport, and capability metadata; no issues found. |
| providers/runtime/factory.py | Wires Cohere into runtime provider
factory dispatch; no issues found. |
| config/settings.py | Adds Cohere API key and proxy settings aliases;
no issues found. |
| api/admin_config/provider_manifest.py | Adds Cohere API key
labeling/description through catalog-derived Admin fields; no issues
found. |
| smoke/lib/config.py | Adds Cohere default smoke model and credential
detection; no issues found. |
| tests/providers/test_cohere.py | Adds request-policy and streaming
adapter tests for the Cohere provider; no issues found. |
| tests/providers/test_provider_runtime.py | Adds Cohere descriptor,
config build, and factory instantiation coverage; no issues found. |
| README.md | Adds Cohere setup instructions and updates provider
counts/order; no issues found. |
| pyproject.toml | Bumps package version for the new provider feature;
no issues found. |
| uv.lock | Updates the lockfile package version to match
`pyproject.toml`; no issues found. |

</details>

<details open><summary><h3>Sequence Diagram</h3></summary>

<a href="#gh-light-mode-only">

```mermaid
%%{init: {'theme': 'neutral'}}%%
sequenceDiagram
participant User as User/Admin config
participant Catalog as Provider catalog/settings
participant Factory as Runtime factory
participant Cohere as CohereProvider
participant Transport as OpenAI chat transport
participant API as Cohere Compatibility API

User->>Catalog: "Configure MODEL=cohere/... and COHERE_API_KEY"
Catalog->>Factory: Build ProviderConfig for provider_id cohere
Factory->>Cohere: Instantiate CohereProvider(config)
Cohere->>Transport: Build chat body with Cohere policy
Transport->>API: "POST /chat/completions stream=true"
API-->>Transport: Streaming OpenAI-compatible chunks
Transport-->>User: Anthropic SSE events
```

</a>
<a href="#gh-dark-mode-only">

```mermaid
%%{init: {'theme': 'base', 'themeVariables': {"darkMode": true, "background": "#0d1117", "primaryColor": "#21262d", "primaryTextColor": "#e6edf3", "primaryBorderColor": "#8b949e", "lineColor": "#8b949e", "textColor": "#e6edf3", "edgeLabelBackground": "#161b22", "actorBkg": "#21262d", "actorBorder": "#8b949e", "actorTextColor": "#e6edf3", "actorLineColor": "#8b949e", "signalColor": "#8b949e", "signalTextColor": "#e6edf3", "noteBkgColor": "#373320", "noteBorderColor": "#d4a72c", "noteTextColor": "#f0e6c0", "labelBoxBkgColor": "#21262d", "labelBoxBorderColor": "#8b949e", "labelTextColor": "#e6edf3", "loopTextColor": "#e6edf3", "activationBkgColor": "#30363d", "activationBorderColor": "#8b949e"}}}%%
sequenceDiagram
participant User as User/Admin config
participant Catalog as Provider catalog/settings
participant Factory as Runtime factory
participant Cohere as CohereProvider
participant Transport as OpenAI chat transport
participant API as Cohere Compatibility API

User->>Catalog: "Configure MODEL=cohere/... and COHERE_API_KEY"
Catalog->>Factory: Build ProviderConfig for provider_id cohere
Factory->>Cohere: Instantiate CohereProvider(config)
Cohere->>Transport: Build chat body with Cohere policy
Transport->>API: "POST /chat/completions stream=true"
API-->>Transport: Streaming OpenAI-compatible chunks
Transport-->>User: Anthropic SSE events
```

</a>
</details>

<sub>Reviews (1): Last reviewed commit: ["Add Cohere
provider"](https://github.com/alishahryar1/free-claude-code/commit/7956802968fc6ce63bb71fe6d7503d484df56b79)
| [Re-trigger
Greptile](https://app.greptile.com/api/retrigger?id=41887611)</sub>

<!-- /greptile_comment -->
2026-07-05 01:01:19 -07:00
Ali Khokhar d4683bf3f6 Add Hugging Face inference provider (#985)
## Problem

FCC did not expose Hugging Face Inference Providers as a selectable
backend. Voice transcription also used the legacy `HF_TOKEN` setting
instead of the canonical Hugging Face API key.

## Changes

| Before | After |
| --- | --- |
| Hugging Face models could not be selected through provider-prefixed
routing. | Hugging Face routes through a thin OpenAI-chat provider using
`huggingface/<model>`. |
| Provider credentials did not include `HUGGINGFACE_API_KEY`. | Admin
config, settings, smoke config, and docs use `HUGGINGFACE_API_KEY`. |
| `HF_TOKEN` remained a voice-only config key. | Owned dotenv files
migrate `HF_TOKEN` to `HUGGINGFACE_API_KEY`, while explicit
`FCC_ENV_FILE` users get a warning. |
| Version metadata stayed on `2.6.0`. | Version metadata moves to
`3.0.0` with a refreshed lockfile. |

<!-- greptile_comment -->

<details open><summary><h3>Greptile Summary</h3></summary>

This PR adds Hugging Face Inference Providers as a selectable backend.
The main changes are:

- Adds a `huggingface` provider using the shared OpenAI-compatible chat
transport.
- Wires `HUGGINGFACE_API_KEY` and `HUGGINGFACE_PROXY` through settings,
Admin UI, provider catalog, runtime factory, and smoke config.
- Migrates owned dotenv files from `HF_TOKEN` to `HUGGINGFACE_API_KEY`
and warns for explicit `FCC_ENV_FILE` users.
- Updates voice transcription plumbing to use the canonical Hugging Face
key.
- Updates docs, examples, version metadata, lockfile, and related tests.
</details>

<h3>Confidence Score: 5/5</h3>

Safe to merge with minimal risk.

No blocking correctness or security issues were identified. The new
provider reuses the existing OpenAI-chat transport pattern. Provider
wiring, env migration, Admin UI, smoke config, voice plumbing, and tests
are consistent.

No files require special attention.

<details><summary><h3><a href="https://www.greptile.com/trex"><img
alt="T-Rex"
src="https://greptile-static-assets.s3.amazonaws.com/trex/trex_green.svg"
height="20" align="absmiddle"></a> T-Rex Logs</h3></summary>

**What T-Rex did**
- The Pytest suite for providers, runtime, env migrations, config, and
contract tests ran and completed with exit code 0 and 198 tests passed.
- The HuggingFace runtime validator script ran and completed
successfully, printing provider\_class=HuggingFaceProvider,
default\_base\_url=https://router.huggingface.co/v1,
credential\_env=HUGGINGFACE\_API\_KEY, and
admin\_field=HUGGINGFACE\_API\_KEY:\[REDACTED\].
- Logs from both runs were captured as artifacts to aid review.

<a
href="https://app.greptile.com/trex/runs/13308618/artifacts"><picture><source
media="(prefers-color-scheme: dark)"
srcset="https://greptile-static-assets.s3.amazonaws.com/badges/ViewAllArtifactsDark.svg?v=4"><source
media="(prefers-color-scheme: light)"
srcset="https://greptile-static-assets.s3.amazonaws.com/badges/ViewAllArtifacts.svg?v=4"><img
alt="View all artifacts"
src="https://greptile-static-assets.s3.amazonaws.com/badges/ViewAllArtifacts.svg?v=4"></picture></a>

<sub><a href="https://www.greptile.com/trex"><img alt="T-Rex"
src="https://greptile-static-assets.s3.amazonaws.com/trex/trex_green.svg"
height="14" align="absmiddle"></a> Ran code and verified through
T-Rex</sub>
</details>

<details open><summary><h3>Important Files Changed</h3></summary>

| Filename | Overview |
|----------|----------|
| providers/huggingface/client.py | Implements Hugging Face via the
shared OpenAI-chat transport with `extra_body` passthrough. |
| config/provider_catalog.py | Registers Hugging Face metadata, default
router URL, credential, proxy, and capabilities. |
| providers/runtime/factory.py | Wires Hugging Face into runtime
provider construction. |
| config/env_migrations.py | Adds safe `HF_TOKEN` to
`HUGGINGFACE_API_KEY` dotenv migration helpers for owned env files. |
| config/settings.py | Adds Hugging Face API key/proxy settings and
removes the legacy `hf_token` setting. |
| api/admin_config/manifest.py | Removes the voice-only `HF_TOKEN` field
and adds Hugging Face smoke model configuration. |
| api/admin_config/provider_manifest.py | Adds Admin UI labeling and
description for `HUGGINGFACE_API_KEY`. |
| messaging/transcription.py | Renames local Whisper token handling to
use the canonical Hugging Face API key. |
| smoke/lib/config.py | Adds Hugging Face smoke-test default model and
credential detection. |
| tests/providers/test_huggingface.py | Adds provider tests for Hugging
Face base URL, request body policy, streaming, and cleanup. |

</details>

<details open><summary><h3>Sequence Diagram</h3></summary>

<a href="#gh-light-mode-only">

```mermaid
%%{init: {'theme': 'neutral'}}%%
sequenceDiagram
participant User as Admin/User
participant Settings as Settings + dotenv migration
participant Catalog as Provider Catalog
participant Runtime as Provider Runtime Factory
participant HF as HuggingFaceProvider
participant Router as router.huggingface.co/v1

User->>Settings: "Configure MODEL=huggingface/<model> and HUGGINGFACE_API_KEY"
Settings->>Settings: Rename owned HF_TOKEN to HUGGINGFACE_API_KEY when present
Settings->>Catalog: Resolve huggingface descriptor and credential/proxy attrs
Catalog->>Runtime: Build ProviderConfig for huggingface
Runtime->>HF: Create HuggingFaceProvider
HF->>Router: Stream OpenAI-compatible chat completion
Router-->>HF: Streaming chunks
HF-->>User: Anthropic SSE response
```

</a>
<a href="#gh-dark-mode-only">

```mermaid
%%{init: {'theme': 'base', 'themeVariables': {"darkMode": true, "background": "#0d1117", "primaryColor": "#21262d", "primaryTextColor": "#e6edf3", "primaryBorderColor": "#8b949e", "lineColor": "#8b949e", "textColor": "#e6edf3", "edgeLabelBackground": "#161b22", "actorBkg": "#21262d", "actorBorder": "#8b949e", "actorTextColor": "#e6edf3", "actorLineColor": "#8b949e", "signalColor": "#8b949e", "signalTextColor": "#e6edf3", "noteBkgColor": "#373320", "noteBorderColor": "#d4a72c", "noteTextColor": "#f0e6c0", "labelBoxBkgColor": "#21262d", "labelBoxBorderColor": "#8b949e", "labelTextColor": "#e6edf3", "loopTextColor": "#e6edf3", "activationBkgColor": "#30363d", "activationBorderColor": "#8b949e"}}}%%
sequenceDiagram
participant User as Admin/User
participant Settings as Settings + dotenv migration
participant Catalog as Provider Catalog
participant Runtime as Provider Runtime Factory
participant HF as HuggingFaceProvider
participant Router as router.huggingface.co/v1

User->>Settings: "Configure MODEL=huggingface/<model> and HUGGINGFACE_API_KEY"
Settings->>Settings: Rename owned HF_TOKEN to HUGGINGFACE_API_KEY when present
Settings->>Catalog: Resolve huggingface descriptor and credential/proxy attrs
Catalog->>Runtime: Build ProviderConfig for huggingface
Runtime->>HF: Create HuggingFaceProvider
HF->>Router: Stream OpenAI-compatible chat completion
Router-->>HF: Streaming chunks
HF-->>User: Anthropic SSE response
```

</a>
</details>

<sub>Reviews (1): Last reviewed commit: ["Add Hugging Face inference
provider"](https://github.com/alishahryar1/free-claude-code/commit/7341d9a923986ac84d5e4fdf858128f913f3e5d3)
| [Re-trigger
Greptile](https://app.greptile.com/api/retrigger?id=41885587)</sub>

<!-- /greptile_comment -->
2026-07-05 00:26:40 -07:00
Ali Khokhar 020bbef64b Add Vercel AI Gateway provider (#984)
## Problem

FCC did not expose Vercel AI Gateway as a provider, so users with
`AI_GATEWAY_API_KEY` could not route Claude, Codex, or messaging
workflows through Vercel's model gateway.

## Changes

| Before | After |
| --- | --- |
| Provider metadata skipped Vercel AI Gateway. | Provider metadata
includes `vercel` with `AI_GATEWAY_API_KEY`, `VERCEL_AI_GATEWAY_PROXY`,
and OpenAI-chat capabilities. |
| No Vercel provider package or factory existed. | `VercelProvider` uses
the shared OpenAI-chat transport with `max_tokens` and preserved
`extra_body`. |
| Admin, docs, smoke config, and model parsing had no Vercel surface. |
Admin, docs, smoke config, and model parsing include Vercel model refs
such as `vercel/openai/gpt-5.5`. |

<!-- greptile_comment -->

<details open><summary><h3>Greptile Summary</h3></summary>

This PR adds Vercel AI Gateway as a new OpenAI-compatible provider. The
main changes are:

- Provider catalog, settings, Admin UI metadata, and runtime factory
wiring for `vercel`.
- A thin `VercelProvider` adapter that reuses the shared OpenAI-chat
streaming transport.
- Vercel-specific docs, environment examples, proxy settings, and
smoke-test defaults.
- Config, contract, runtime, and provider tests covering the new
provider path.
- Package version and lockfile updates for the new feature.
</details>

<h3>Confidence Score: 5/5</h3>

This PR is safe to merge with minimal risk.

The new provider follows the existing catalog, settings, factory, and
shared transport patterns. The change includes focused config, runtime,
smoke, and provider tests. The package version and lockfile were updated
with the production changes. No functional or security issues were
identified in the changed paths.

No files require special attention.

<details><summary><h3><a href="https://www.greptile.com/trex"><img
alt="T-Rex"
src="https://greptile-static-assets.s3.amazonaws.com/trex/trex_green.svg"
height="20" align="absmiddle"></a> T-Rex Logs</h3></summary>

**What T-Rex did**
- A focused pytest run for the Vercel provider completed successfully
with 187 tests passed in 3.93 seconds and EXIT\_CODE: 0.
- An offline probe script named vercel-provider-offline-probe.py was
generated to exercise the real factory/provider/request construction
code offline.
- The offline probe log vercel-provider-offline-probe.log showed the
expected provider setup and a successful exit, including
catalog\_has\_vercel=True, factory\_has\_vercel=True, provider class
VercelProvider, base URL, synthetic API key propagation, max\_tokens
preserved, max\_completion\_tokens absent, extra\_body preserved, and
EXIT\_CODE: 0.

<a
href="https://app.greptile.com/trex/runs/13307261/artifacts"><picture><source
media="(prefers-color-scheme: dark)"
srcset="https://greptile-static-assets.s3.amazonaws.com/badges/ViewAllArtifactsDark.svg?v=4"><source
media="(prefers-color-scheme: light)"
srcset="https://greptile-static-assets.s3.amazonaws.com/badges/ViewAllArtifacts.svg?v=4"><img
alt="View all artifacts"
src="https://greptile-static-assets.s3.amazonaws.com/badges/ViewAllArtifacts.svg?v=4"></picture></a>

<sub><a href="https://www.greptile.com/trex"><img alt="T-Rex"
src="https://greptile-static-assets.s3.amazonaws.com/trex/trex_green.svg"
height="14" align="absmiddle"></a> Ran code and verified through
T-Rex</sub>
</details>

<details open><summary><h3>Important Files Changed</h3></summary>

| Filename | Overview |
|----------|----------|
| README.md | Adds Vercel AI Gateway setup guidance and renumbers
provider documentation. |
| api/admin_config/provider_manifest.py | Adds Admin UI field metadata
for `AI_GATEWAY_API_KEY` via existing catalog-derived manifest flow. |
| config/provider_catalog.py | Registers `vercel` as an OpenAI-chat
provider with gateway credential, default base URL, proxy, and
capabilities. |
| config/settings.py | Adds settings fields for Vercel gateway API key
and proxy aliases. |
| providers/runtime/factory.py | Wires the new `vercel` provider id to
`VercelProvider` in runtime factory registration. |
| providers/vercel/client.py | Implements a thin Vercel adapter over
shared OpenAI-chat transport with `max_tokens` and `extra_body`
passthrough. |
| pyproject.toml | Bumps the package version to `2.6.0` for the new
provider feature. |
| smoke/lib/config.py | Adds Vercel smoke defaults and credential
detection for provider smoke selection. |
| tests/providers/test_provider_runtime.py | Adds runtime config and
factory instantiation coverage for the Vercel provider. |
| tests/providers/test_vercel.py | Adds unit tests for Vercel base URL
handling, request-body policy, streaming deltas, and cleanup. |
| uv.lock | Synchronizes the lockfile package version with
`pyproject.toml`. |

</details>

<details open><summary><h3>Sequence Diagram</h3></summary>

<a href="#gh-light-mode-only">

```mermaid
%%{init: {'theme': 'neutral'}}%%
sequenceDiagram
participant User as User/Admin config
participant Settings as Settings/env
participant Catalog as Provider catalog
participant Factory as Runtime factory
participant Vercel as VercelProvider
participant Gateway as Vercel AI Gateway

User->>Settings: "Set AI_GATEWAY_API_KEY and MODEL=vercel/..."
Settings->>Catalog: Resolve vercel descriptor
Catalog->>Factory: Build ProviderConfig with key/base/proxy
Factory->>Vercel: Instantiate VercelProvider
Vercel->>Gateway: Stream OpenAI Chat Completions
Gateway-->>Vercel: OpenAI-compatible chunks
Vercel-->>User: Anthropic SSE via shared transport
```

</a>
<a href="#gh-dark-mode-only">

```mermaid
%%{init: {'theme': 'base', 'themeVariables': {"darkMode": true, "background": "#0d1117", "primaryColor": "#21262d", "primaryTextColor": "#e6edf3", "primaryBorderColor": "#8b949e", "lineColor": "#8b949e", "textColor": "#e6edf3", "edgeLabelBackground": "#161b22", "actorBkg": "#21262d", "actorBorder": "#8b949e", "actorTextColor": "#e6edf3", "actorLineColor": "#8b949e", "signalColor": "#8b949e", "signalTextColor": "#e6edf3", "noteBkgColor": "#373320", "noteBorderColor": "#d4a72c", "noteTextColor": "#f0e6c0", "labelBoxBkgColor": "#21262d", "labelBoxBorderColor": "#8b949e", "labelTextColor": "#e6edf3", "loopTextColor": "#e6edf3", "activationBkgColor": "#30363d", "activationBorderColor": "#8b949e"}}}%%
sequenceDiagram
participant User as User/Admin config
participant Settings as Settings/env
participant Catalog as Provider catalog
participant Factory as Runtime factory
participant Vercel as VercelProvider
participant Gateway as Vercel AI Gateway

User->>Settings: "Set AI_GATEWAY_API_KEY and MODEL=vercel/..."
Settings->>Catalog: Resolve vercel descriptor
Catalog->>Factory: Build ProviderConfig with key/base/proxy
Factory->>Vercel: Instantiate VercelProvider
Vercel->>Gateway: Stream OpenAI Chat Completions
Gateway-->>Vercel: OpenAI-compatible chunks
Vercel-->>User: Anthropic SSE via shared transport
```

</a>
</details>

<sub>Reviews (1): Last reviewed commit: ["Add Vercel AI Gateway
provider"](https://github.com/alishahryar1/free-claude-code/commit/862ae946b2f7029ce4d3b309178312c9afe2751b)
| [Re-trigger
Greptile](https://app.greptile.com/api/retrigger?id=41883209)</sub>

<!-- /greptile_comment -->
2026-07-04 23:49:48 -07:00
Ali Khokhar 85b601884d Remove legacy future annotation imports (#982)
## Problem

Python 3.14 provides native lazy annotations, but the codebase still
relied on legacy future annotation imports. Those imports also made
type-only import cycles easier to hide instead of fixing ownership
boundaries.

## Changes

| Before | After |
| --- | --- |
| Python files used `from __future__ import annotations`. | Python files
rely on Python 3.14 native lazy annotations. |
| Some runtime modules used `TYPE_CHECKING` or local imports for
required dependencies. | Runtime modules use top-level owner-module
imports with explicit boundaries. |
| Local and GitHub guardrails only rejected type ignore suppressions. |
Local and GitHub guardrails reject type ignore suppressions and legacy
future annotation imports. |
| Agent docs only documented the no-type-ignore rule. | Agent docs
document the Python 3.14 annotation and import-boundary rules. |

<!-- greptile_comment -->

<details open><summary><h3>Greptile Summary</h3></summary>

This PR moves the codebase to Python 3.14 native lazy annotations. The
main changes are:

- Removed legacy `from __future__ import annotations` imports across
Python modules.
- Promoted selected runtime dependencies from `TYPE_CHECKING` or local
imports to explicit owner-module imports.
- Added local, GitHub, and contract-test guardrails to reject legacy
future annotation imports.
- Updated agent docs with the annotation and import-boundary rules.
- Bumped the package patch version for production-file changes.
</details>

<h3>Confidence Score: 5/5</h3>

Safe to merge with low risk.

The changes are mostly mechanical annotation cleanup with matching CI
and contract-test guardrails. Reviewed import-boundary updates did not
show a confirmed runtime cycle or dependency break.

No files require special attention.

<details><summary><h3><a href="https://www.greptile.com/trex"><img
alt="T-Rex"
src="https://greptile-static-assets.s3.amazonaws.com/trex/trex_green.svg"
height="20" align="absmiddle"></a> T-Rex Logs</h3></summary>

**What T-Rex did**
- Performed an end-to-end validation of the guardrail contract suite: an
environment check confirmed uv availability, a guardrail pytest run used
CPython 3.14.0 with 5 passing contract tests, 3 focused CI-script tests
passed, and the direct CI suppressions guardrail command (including the
legacy future-annotations grep) also passed.

<a
href="https://app.greptile.com/trex/runs/13303335/artifacts"><picture><source
media="(prefers-color-scheme: dark)"
srcset="https://greptile-static-assets.s3.amazonaws.com/badges/ViewAllArtifactsDark.svg?v=4"><source
media="(prefers-color-scheme: light)"
srcset="https://greptile-static-assets.s3.amazonaws.com/badges/ViewAllArtifacts.svg?v=4"><img
alt="View all artifacts"
src="https://greptile-static-assets.s3.amazonaws.com/badges/ViewAllArtifacts.svg?v=4"></picture></a>

<sub><a href="https://www.greptile.com/trex"><img alt="T-Rex"
src="https://greptile-static-assets.s3.amazonaws.com/trex/trex_green.svg"
height="14" align="absmiddle"></a> Ran code and verified through
T-Rex</sub>
</details>

<details open><summary><h3>Important Files Changed</h3></summary>

| Filename | Overview |
|----------|----------|
| api/runtime.py | Moves messaging, CLI manager, session, limiter, and
tree dependencies from local/type-checking imports to explicit top-level
owner-module imports. |
| messaging/platforms/telegram.py | Removes future annotations and
promotes Telegram SDK type imports into the existing availability guard.
|
| messaging/platforms/telegram_inbound.py | Removes future annotations
and imports Telegram SDK types at module scope for inbound
normalization. |
| tests/contracts/test_import_boundaries.py | Adds an AST contract that
rejects legacy future annotation imports across Python files. |
| scripts/ci.sh | Extends the local suppression check to reject legacy
future annotation imports alongside type-ignore suppressions. |
| scripts/ci.ps1 | Mirrors the local PowerShell CI suppression check for
legacy future annotations. |
| .github/workflows/tests.yml | Renames and broadens the GitHub
guardrail job to reject both type suppressions and legacy future
annotations. |
| pyproject.toml | Bumps the patch version for production-file changes.
|

</details>

<details open><summary><h3>Sequence Diagram</h3></summary>

<a href="#gh-light-mode-only">

```mermaid
%%{init: {'theme': 'neutral'}}%%
sequenceDiagram
participant Dev as Developer/CI
participant Guard as Suppression guard
participant AST as Import-boundary contract test
participant Py as Python modules

Dev->>Guard: Run local/GitHub suppression check
Guard->>Py: "Scan *.py for type ignores and future annotations"
Guard-->>Dev: Fail if legacy annotation import remains
Dev->>AST: Run pytest contract tests
AST->>Py: Parse imports with ast
AST-->>Dev: Assert no future annotations/import-boundary violations
Py-->>Dev: Use Python 3.14 native lazy annotations
```

</a>
<a href="#gh-dark-mode-only">

```mermaid
%%{init: {'theme': 'base', 'themeVariables': {"darkMode": true, "background": "#0d1117", "primaryColor": "#21262d", "primaryTextColor": "#e6edf3", "primaryBorderColor": "#8b949e", "lineColor": "#8b949e", "textColor": "#e6edf3", "edgeLabelBackground": "#161b22", "actorBkg": "#21262d", "actorBorder": "#8b949e", "actorTextColor": "#e6edf3", "actorLineColor": "#8b949e", "signalColor": "#8b949e", "signalTextColor": "#e6edf3", "noteBkgColor": "#373320", "noteBorderColor": "#d4a72c", "noteTextColor": "#f0e6c0", "labelBoxBkgColor": "#21262d", "labelBoxBorderColor": "#8b949e", "labelTextColor": "#e6edf3", "loopTextColor": "#e6edf3", "activationBkgColor": "#30363d", "activationBorderColor": "#8b949e"}}}%%
sequenceDiagram
participant Dev as Developer/CI
participant Guard as Suppression guard
participant AST as Import-boundary contract test
participant Py as Python modules

Dev->>Guard: Run local/GitHub suppression check
Guard->>Py: "Scan *.py for type ignores and future annotations"
Guard-->>Dev: Fail if legacy annotation import remains
Dev->>AST: Run pytest contract tests
AST->>Py: Parse imports with ast
AST-->>Dev: Assert no future annotations/import-boundary violations
Py-->>Dev: Use Python 3.14 native lazy annotations
```

</a>
</details>

<sub>Reviews (2): Last reviewed commit: ["Remove legacy future
annotations
import"](https://github.com/alishahryar1/free-claude-code/commit/6e6cda69da243bbdb92831207aecb3731ad469f8)
| [Re-trigger
Greptile](https://app.greptile.com/api/retrigger?id=41875785)</sub>

<!-- /greptile_comment -->
2026-07-04 21:41:51 -07:00