Repository navigation
Conversation
|
Thanks for your pull request! It looks like this may be your first contribution to a Google open source project. Before we can look at your pull request, you'll need to sign a Contributor License Agreement (CLA). View this failed invocation of the CLA check for more information. For the most up to date status, view the checks section at the bottom of the pull request. |
The streaming path rebuilt each tool call from its id, name and arguments only, so the thought_signature Gemini 3 attaches to a call (in extra_content on the Vertex AI OpenAI-compatible endpoint) never reached the function-call part. The next request then failed with "Function call is missing a thought_signature in functionCall parts". The non-streaming path kept it. Carry the signature on FunctionChunk, keep it per tool-call index while the stream is assembled, and set it on the assembled tool call, so the existing extraction restores it onto the part as it does for non-streamed responses. Fixes google#7438
vetler
force-pushed
the
fix/litellm-streamed-thought-signature
branch
from
October 7, 2026 17:01
678a42e to
42a5da3
Compare
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Link to Issue or Description of Change
1. Link to an existing issue (if applicable):
Problem:
With
LiteLlmandstream=True, thethought_signatureGemini 3 attaches to a tool call is dropped while the streamed call is reassembled.FunctionChunkonly carriedid,name,argsandindex, and the assembledChatCompletionMessageToolCallwas built from those, so_extract_thought_signature_from_tool_callfound nothing. The stored function-call part had no signature and Gemini 3 rejected the next request:The non-streaming path keeps the signature (#3627, #4650).
Solution:
FunctionChunkgets an optionalthought_signature. Gemini signs only the first call of a parallel batch, so most chunks carry none._model_response_to_chunkfills it with the existing_extract_thought_signature_from_tool_call, so all the places a signature can arrive are covered (extra_content.google,provider_specific_fields, an id with an embedded signature)._finalize_tool_call_responsesets it asextra_content.google.thought_signatureon the assembled tool call._message_to_generate_content_responsethen restores it onto the part exactly as for a non-streamed response, so no second code path converts signatures.Chunks with neither a name nor arguments are still skipped as before, so no call is created from a chunk that only carries a signature. LiteLLM's Gemini provider and the Vertex AI OpenAI-compatible endpoint both send the signature on the delta that carries the call's name.
Testing Plan
Unit Tests:
New tests in
tests/unittests/models/test_litellm.py:test_model_response_to_chunk_keeps_thought_signature: the signature on a streamed delta reaches theFunctionChunk.test_streaming_tool_call_keeps_thought_signature: the signature on the first of two argument fragments survives assembly, and_content_to_message_paramsends it back on the follow-up request.test_streaming_parallel_tool_calls_keep_signature_per_call: with two parallel calls where only the first is signed, the first part keeps its signature and the second staysNone.All three fail on
mainand pass with the fix.Notes on that run:
kuberneteswas pinned to 36.0.3, the version in the constraints files. 37.0.0 (released 2026-10-07) makes fivetest_gke_code_executor.pytests fail onmaintoo, andmain's CI fails on the same five.test_session_context.py::test_timeout_during_connection: a wall-clock check (start()took 1.9s against a 1.0s limit) while four suites ran at once. It passed when py312 ran on its own.sitecustomizemodule thattest_import_loading.pyflags.pre-commit run --fileson both changed files passes, and mypy reports no new errors inlite_llm.py.Manual End-to-End (E2E) Tests:
The reproduction script from LiteLlm streaming drops Gemini thought_signature from tool calls; Gemini 3 then rejects the follow-up request #7438 (LiteLLM client mocked with the shape the Vertex AI OpenAI-compatible endpoint returns):
Live, against
gemini-3-flash-previewon the Vertex AI OpenAI-compatible endpoint: aRunnerwithRunConfig(streaming_mode=StreamingMode.SSE)and anLlmAgentwith one tool, authenticated with Application Default Credentials.Script (
VERTEX_PROJECT=<project> python e2e_live.py)On
main, the stored call has no signature and Vertex rejects the follow-up request (traceback trimmed):With this change the signature is kept and the turn completes:
Checklist
Additional context
Gemini 3 checks signatures only for function calls in the current turn, so the bug can go unnoticed: if a user-role message follows the function response (for example the
instructionADK adds as user content whenstatic_instructionis set), the request succeeds without the signature.🤖 Generated with Claude Code