Skip to content

fix(litellm): keep Gemini thought signatures on streamed tool calls - #7441

Open
vetler wants to merge 2 commits into
google:mainfrom
vetler:fix/litellm-streamed-thought-signature
Open

vetler wants to merge 2 commits into
google:mainfrom
vetler:fix/litellm-streamed-thought-signature

Conversation

@vetler

@vetler vetler commented Oct 7, 2026

Copy link
Copy Markdown

Link to Issue or Description of Change

1. Link to an existing issue (if applicable):

Problem:

With LiteLlm and stream=True, the thought_signature Gemini 3 attaches to a tool call is dropped while the streamed call is reassembled. FunctionChunk only carried id, name, args and index, and the assembled ChatCompletionMessageToolCall was built from those, so _extract_thought_signature_from_tool_call found nothing. The stored function-call part had no signature and Gemini 3 rejected the next request:

Function call is missing a thought_signature in functionCall parts.

The non-streaming path keeps the signature (#3627, #4650).

Solution:

  • FunctionChunk gets an optional thought_signature. Gemini signs only the first call of a parallel batch, so most chunks carry none.
  • _model_response_to_chunk fills it with the existing _extract_thought_signature_from_tool_call, so all the places a signature can arrive are covered (extra_content.google, provider_specific_fields, an id with an embedded signature).
  • The streaming loop keeps it per tool-call index, next to name, args and id.
  • _finalize_tool_call_response sets it as extra_content.google.thought_signature on the assembled tool call. _message_to_generate_content_response then restores it onto the part exactly as for a non-streamed response, so no second code path converts signatures.

Chunks with neither a name nor arguments are still skipped as before, so no call is created from a chunk that only carries a signature. LiteLLM's Gemini provider and the Vertex AI OpenAI-compatible endpoint both send the signature on the delta that carries the call's name.

Testing Plan

Unit Tests:

  • I have added or updated unit tests for my change.
  • All unit tests pass locally.

New tests in tests/unittests/models/test_litellm.py:

  • test_model_response_to_chunk_keeps_thought_signature: the signature on a streamed delta reaches the FunctionChunk.
  • test_streaming_tool_call_keeps_thought_signature: the signature on the first of two argument fragments survives assembly, and _content_to_message_param sends it back on the follow-up request.
  • test_streaming_parallel_tool_calls_keep_signature_per_call: with two parallel calls where only the first is signed, the first part keeps its signature and the second stays None.

All three fail on main and pass with the fix.

$ pytest tests/unittests/models/ tests/unittests/flows -q
2502 passed in 55.18s

$ tox -p auto   # pytest tests/unittests on Python 3.11-3.14
py311: OK
py312: OK   (17912 passed, rerun on its own; see below)
py313: OK
py314: OK

Notes on that run:

  • kubernetes was pinned to 36.0.3, the version in the constraints files. 37.0.0 (released 2026-10-07) makes five test_gke_code_executor.py tests fail on main too, and main's CI fails on the same five.
  • In the parallel run, py312 had one failure, test_session_context.py::test_timeout_during_connection: a wall-clock check (start() took 1.9s against a 1.0s limit) while four suites ran at once. It passed when py312 ran on its own.
  • The environments used uv's managed Python. Homebrew's Python ships a sitecustomize module that test_import_loading.py flags.

pre-commit run --files on both changed files passes, and mypy reports no new errors in lite_llm.py.

Manual End-to-End (E2E) Tests:

  1. The reproduction script from LiteLlm streaming drops Gemini thought_signature from tool calls; Gemini 3 then rejects the follow-up request #7438 (LiteLLM client mocked with the shape the Vertex AI OpenAI-compatible endpoint returns):

    # main
    non-streamed thought_signature: b'opaque-signature'
    streamed thought_signature:     None
    
    # this branch
    non-streamed thought_signature: b'opaque-signature'
    streamed thought_signature:     b'opaque-signature'
    
  2. Live, against gemini-3-flash-preview on the Vertex AI OpenAI-compatible endpoint: a Runner with RunConfig(streaming_mode=StreamingMode.SSE) and an LlmAgent with one tool, authenticated with Application Default Credentials.

    Script (VERTEX_PROJECT=<project> python e2e_live.py)
    """Live Runner E2E for google/adk-python#7438 against Gemini 3 on Vertex AI."""
    
    import asyncio
    import os
    import sys
    
    import google.auth
    import google.auth.transport.requests
    from google.adk.agents import LlmAgent
    from google.adk.agents.run_config import RunConfig
    from google.adk.agents.run_config import StreamingMode
    from google.adk.models.lite_llm import LiteLlm
    from google.adk.runners import Runner
    from google.adk.sessions import InMemorySessionService
    from google.genai import types
    
    PROJECT = os.environ["VERTEX_PROJECT"]
    MODEL = os.environ.get("VERTEX_MODEL", "google/gemini-3-flash-preview")
    API_BASE = (
        f"https://aiplatform.googleapis.com/v1/projects/{PROJECT}"
        "/locations/global/endpoints/openapi"
    )
    
    
    def get_weather(city: str) -> dict:
      """Returns the weather for a city."""
      return {"city": city, "forecast": "rain", "temperature_c": 9}
    
    
    async def main():
      credentials, _ = google.auth.default(
          scopes=["https://www.googleapis.com/auth/cloud-platform"]
      )
      credentials.refresh(google.auth.transport.requests.Request())
      agent = LlmAgent(
          name="weather",
          model=LiteLlm(
              model=MODEL,
              api_base=API_BASE,
              api_key=credentials.token,
              custom_llm_provider="openai",
          ),
          instruction="Answer weather questions. Always call get_weather first.",
          tools=[get_weather],
      )
      sessions = InMemorySessionService()
      runner = Runner(app_name="e2e", agent=agent, session_service=sessions)
      session = await sessions.create_session(app_name="e2e", user_id="u")
      try:
        async for event in runner.run_async(
            user_id="u",
            session_id=session.id,
            new_message=types.Content(
                role="user", parts=[types.Part(text="What's the weather in Oslo?")]
            ),
            run_config=RunConfig(streaming_mode=StreamingMode.SSE),
        ):
          if event.partial:
            continue
          if event.error_code or event.error_message:
            print(f"[{event.author}] error {event.error_code}: {event.error_message}")
          for part in event.content.parts if event.content else []:
            if part.function_call:
              sig = part.thought_signature
              print(
                  f"[{event.author}] function_call {part.function_call.name}"
                  f"({part.function_call.args}) thought_signature="
                  + (f"<{len(sig)} bytes>" if sig else "None")
              )
            elif part.function_response:
              print(f"[{event.author}] function_response {part.function_response.name}")
            elif part.text and not part.thought:
              print(f"[{event.author}] text: {part.text.strip()}")
      except Exception as e:  # pylint: disable=broad-exception-caught
        print(f"{type(e).__name__}: {str(e)[:400]}")
        sys.exit(1)
    
    
    asyncio.run(main())

    On main, the stored call has no signature and Vertex rejects the follow-up request (traceback trimmed):

    [weather] function_call get_weather({'city': 'Oslo'}) thought_signature=None
    [weather] function_response get_weather
    BadRequestError: litellm.BadRequestError: OpenAIException - Error code: 400 - [{'error': {'code': 400, 'message': 'Unable to submit request because function call `get_weather` in the 2. content block is missing a `thought_signature`. Learn more: https://docs.cloud.google.com/vertex-ai/generative-ai/docs/thought-signatures', 'status': 'INVALID_ARGUMENT'}}]
    

    With this change the signature is kept and the turn completes:

    [weather] function_call get_weather({'city': 'Oslo'}) thought_signature=<355 bytes>
    [weather] function_response get_weather
    [weather] text: The weather in Oslo is currently 9°C with rain.
    

Checklist

  • I have read the CONTRIBUTING.md document.
  • I have performed a self-review of my own code.
  • I have commented my code, particularly in hard-to-understand areas.
  • I have added tests that prove my fix is effective or that my feature works.
  • New and existing unit tests pass locally with my changes.
  • I have manually tested my changes end-to-end.
  • Any dependent changes have been merged and published in downstream modules.

Additional context

Gemini 3 checks signatures only for function calls in the current turn, so the bug can go unnoticed: if a user-role message follows the function response (for example the instruction ADK adds as user content when static_instruction is set), the request succeeds without the signature.

🤖 Generated with Claude Code

@google-cla

google-cla Bot commented Oct 7, 2026

Copy link
Copy Markdown

Thanks for your pull request! It looks like this may be your first contribution to a Google open source project. Before we can look at your pull request, you'll need to sign a Contributor License Agreement (CLA).

View this failed invocation of the CLA check for more information.

For the most up to date status, view the checks section at the bottom of the pull request.

The streaming path rebuilt each tool call from its id, name and arguments
only, so the thought_signature Gemini 3 attaches to a call (in extra_content
on the Vertex AI OpenAI-compatible endpoint) never reached the function-call
part. The next request then failed with "Function call is missing a
thought_signature in functionCall parts". The non-streaming path kept it.

Carry the signature on FunctionChunk, keep it per tool-call index while the
stream is assembled, and set it on the assembled tool call, so the existing
extraction restores it onto the part as it does for non-streamed responses.

Fixes google#7438
@vetler
vetler force-pushed the fix/litellm-streamed-thought-signature branch from 678a42e to 42a5da3 Compare October 7, 2026 17:01
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

LiteLlm streaming drops Gemini thought_signature from tool calls; Gemini 3 then rejects the follow-up request

2 participants