Skip to main content

Try Gemini 3.8 Flash in the Workbench

Run this model interactively, tune parameters, and compare outputs.
Model ID: gemini-3-8-flash gemini-3.8-flash is a Multimodal LLM. It is Google’s Flash-tier model for long-horizon software engineering, autonomous agents, and multi-step enterprise workflows, with gains over Gemini 3.7 Flash across software engineering, agentic tasks, and specialized-domain reasoning. Google reports 54.9% on HLE-Verified and says it outperforms most larger models on DeepSWE v1.1, its long-horizon engineering benchmark. Some other noteworthy features of gemini-3.8-flash include configurable thinking levels (low, medium, high; minimal is not supported and returns an error), structured output, function calling, code execution, context caching, search grounding, URL context, file search, computer use (preview), and batch processing. It accepts text, images, audio, video, and PDFs and outputs text only. On hard tasks it tends to run extra reasoning steps and iterative tool calls, so token usage per task can be higher than 3.7 Flash, especially at higher thinking levels. The Live API, image generation, and audio generation are not supported.

Example request

Use the Workbench as a request builder: configure parameters for this model in the UI, then open the API tab to copy the exact cURL or Python call.

Fetch model details

The models endpoint returns the full model object, including its json_request_schema.

Request parameters

This model follows the standard OpenAI chat completions request body. See the chat completions reference for the full parameter list.