financetom
Technology
financetom
/
Technology
/
Google AI can now focus on individual speakers in a crowd
News World Market Environment Technology Personal Finance Politics Retail Business Economy Cryptocurrency Forex Stocks Market Commodities
Google AI can now focus on individual speakers in a crowd
Apr 13, 2018 3:19 AM

Just as most smartphone cameras now allow users to focus on a single object among many, it may soon be possible to pick out individual voices in a crowd by suppressing all other sounds, thanks to a new Artificial Intelligence (AI) system developed by Google researchers.

This is an important development as computers are not as good as humans at focusing their attention on a particular person in a noisy environment.

Known as the cocktail party effect, the capability to mentally "mute" all other voices and sounds comes natural to us humans.

However, automatic speech separation -- separating an audio signal into its individual speech sources -- remains a significant challenge for computers, Inbar Mosseri and Oran Lang, software engineers at Google Research, wrote in a blog post this week.

In a new paper, the researchers presented a deep learning audio-visual model for isolating a single speech signal from a mixture of sounds such as other voices and background noise.

"In this work, we are able to computationally produce videos in which speech of specific people is enhanced while all other sounds are suppressed," Mosseri and Lang said.

The method works on ordinary videos with a single audio track, and all that is required from the user is to select the face of the person in the video they want to hear, or to have such a person be selected algorithmically based on context.

The researchers believe this capability can have a wide range of applications, from speech enhancement and recognition in videos, through video conferencing, to improved hearing aids, especially in situations where there are multiple people speaking.

"A unique aspect of our technique is in combining both the auditory and visual signals of an input video to separate the speech," the researchers said.

"Intuitively, movements of a person's mouth, for example, should correlate with the sounds produced as that person is speaking, which in turn can help identify which parts of the audio correspond to that person," they explained.

The visual signal not only improves the speech separation quality significantly in cases of mixed speech, but, importantly, it also associates the separated, clean speech tracks with the visible speakers in the video, the researchers said.

First Published:Apr 13, 2018 12:19 PM IST

Comments
Welcome to financetom comments! Please keep conversations courteous and on-topic. To fosterproductive and respectful conversations, you may see comments from our Community Managers.
Sign up to post
Sort by
Show More Comments
Related Articles >
Fortude recognized as Microsoft ‘Fabric Featured Partner’, reinforcing expertise in data analytics
Fortude recognized as Microsoft ‘Fabric Featured Partner’, reinforcing expertise in data analytics
Jul 1, 2026
New York, July 01, 2026 (GLOBE NEWSWIRE) -- Fortude, a global digital solutions company, has been recognized as a Microsoft Fabric Featured Partner, a designation verified by Microsoft engineering and awarded to a select few partners with a proven track record of successful Fabric implementations. This builds on Fortude’s Microsoft Analytics on Azure specialization achieved last September, reinforcing its position...
New Era Energy & Digital Announces Leadership Transition to Support Next Phase of Execution and Growth
New Era Energy & Digital Announces Leadership Transition to Support Next Phase of Execution and Growth
Jul 1, 2026
MIDLAND, Texas, July 01, 2026 (GLOBE NEWSWIRE) -- New Era Energy & Digital, Inc. ( NUAI ) (“New Era” or the “Company”), a developer of next-generation digital infrastructure and integrated power assets, today announced a leadership transition designed to support the Company’s next phase of execution, delivery and growth. Effective July 1, 2026, Charlie  Nelson, currently President and Chief Operating...
Dawnguard launches platform to build secure cloud systems from day zero, with fresh funding and US office
Dawnguard launches platform to build secure cloud systems from day zero, with fresh funding and US office
Jul 1, 2026
New York and Amsterdam, July 01, 2026 (GLOBE NEWSWIRE) -- As AI-assisted engineering accelerates how quickly software is designed, written, and shipped, cybersecurity teams are facing a harder problem: risk is being created earlier than traditional tools can see it. Dawnguard announced the public launch of its security architecture automation platform, making it available to organizations looking to design, build,...
Revvity Expands Signals AI Ecosystem Through Anthropic Claude Integration
Revvity Expands Signals AI Ecosystem Through Anthropic Claude Integration
Jul 1, 2026
New MCP connector extends Signals AI beyond the Signals One platform, enabling scientists to interact with connected R&D knowledge through Claude Claude accesses scientific data and organizational knowledge through Signals' intelligence layer, improving context-aware analysis and decision-making Collaboration combines industry-leading AI capabilities with trusted scientific workflows to help accelerate research and development WALTHAM, Mass.--(BUSINESS WIRE)-- Revvity, Inc. ( RVTY )...
Copyright 2023-2026 - www.financetom.com All Rights Reserved