This week brought sharper voice models from Google, faster video avatars from Tavus, a unified creative connector from ElevenLabs, and a long overdue editing overhaul from Midjourney. Here is what happened this week.
01 · Foundation Models
Google · Gemini 3.8 Live · September 2026
On September 15, 2026, Google released two real time voice models designed for different
use cases. Gemini 3.8 Live is built for fast, cost efficient spoken dialogue, while Gemini 3.8
Live Extended Thinking is designed for complex, multi step tasks that require deeper
reasoning before responding. Both models can run tool and software calls in the background
while keeping a conversation going, and both support 97 languages. The Extended Thinking
variant now holds the top position on Artificial Analysis' Speech to Speech Quality Index, an
independent ranking of conversational AI models. Pricing is set at $0.005 per minute for
audio input and $0.018 per minute for audio output.
Our Takeaway: Developers building voice powered products now have a production
ready model that handles interruptions, language switching, and background task
execution without disrupting a live call. Teams already on the previous Gemini Live model
can upgrade without migration costs, making the move accessible without a budget fight
against comparable offerings from OpenAI.
02 · AI Video
Tavus · Phoenix 4.5 · September 2026
Tavus released Phoenix 4.5 on September 14, 2026, an upgrade to its real time AI human
rendering model. Unlike its predecessor, Phoenix 4.5 extends lifelike motion beyond the face
to include the head, shoulders, posture, and torso. The model responds in 134 milliseconds
from audio input to video output, which Tavus states is 25 percent faster than any competing
model currently available. Phoenix 4.5 is accessible immediately to all users through the PAL
Maker tool and the Tavus developer access program.
Our Takeaway: Product teams and SaaS companies building AI powered onboarding
guides, virtual sales reps, or customer support agents can now deploy video characters
that move and react the way a real person does on a video call. That visual realism reduces
the gap between an automated experience and a human one without requiring additional
development work from the teams deploying it.
03 · Creative AI
ElevenLabs · MCP Connector · September 2026
ElevenLabs updated its hosted connector on September 14, 2026, adding the ability to
generate speech, music, sound effects, images, and video without leaving an active chat
session in AI assistants such as Claude, ChatGPT, and Cursor. Before this update, the
connector handled only agent management tasks. The expanded tool draws on more than 50
generation models and requires only a one time sign in, with no additional configuration or
server setup needed.
Our Takeaway: Marketing and content teams that already use AI assistants to draft
scripts can now take a brief all the way to a finished video with voiceover, music, and
sound design inside a single conversation thread. Removing the multi app assembly work
that previously sat between ideation and final output compresses a process that often
consumed hours into a single session.
04 · AI Image
Midjourney · V8.2 Edit Model · August 2026
Midjourney opened its V8.2 image edit model to all users for testing on August 27, 2026. The
new system merges four tools that previously ran separately: editing an image using plain
written instructions, generating images from up to four reference photos at once, painting
over selected areas of an image, and extending the canvas beyond its original edges. Before
this release, Midjourney's editing tools ran on an older V6.1 foundation even as image
generation had moved to V8, requiring users to navigate two or three different model
versions within a single workflow. The V8.2 edit model consolidates all of this under one
roof.
Our Takeaway: Creative teams and brand designers who rely on Midjourney for visual
production can now move an image through an entire revision cycle, from swapping a
background to expanding the frame, without switching tools or model versions. Collapsing
what was a multi step, multi tool process into a single session frees up time that was
previously spent managing workarounds rather than producing work.
This week's releases share a common direction: tools that previously required multiple
sessions, multiple applications, or multiple versions are being consolidated into single,
continuous workflows. Whether that means a voice call that handles tasks in the background,
an avatar that moves like a full human, a creative suite accessible from one chat window, or
an image editor with no version gaps, the pattern is compression of friction. For business
teams, the practical effect is fewer handoffs between tools and more time spent on the work
itself.
ASR AI Studio