Skip to main content

AI software can identify objects in photos and videos at near-human levels

A new AI software program developed by researchers at Google and Stanford University can recognise objects in photos and videos at near-human levels of understanding.

ai software program google stanford university object recognition technology images videos human level understanding

It was only recently that computer systems became smart enough to identify unknown objects in photographs. Even then, it has generally been limited to individual objects. Now, two separate teams of researchers at Google and Stanford University have created software able to describe entire scenes. This could lead to much better and more intelligent algorithms in the future.
Stanford's work, entitled "Deep Visual-Semantic Alignments for Generating Image Descriptions", explains how specific details found in photographs and videos can be translated into written text. Google's version of the technology, in a study titled "Show and Tell: A Neural Image Caption Generator", produced similar results.
While each team used a slightly different approach, they both combined deep convolutional neural networks with recurrent neural networks that excel at text analysis and natural language processing. The programs were able to "learn" from each new interaction, with algorithms enabling the system to improve its accuracy by scanning scene after scene, looking for patterns, and then using the accumulation of previously described scenes to extrapolate what is being depicted in the next unknown image.

ai image recognition

"The system can analyse an unknown image and explain it in words and phrases that make sense," says Fei-Fei Li, a professor of computer science and director of the Stanford Artificial Intelligence Lab. "This is an important milestone. It's the first time we've had a computer vision system that could tell a basic story about an unknown image by identifying discrete objects and also putting them into some context."
These latest algorithms are being trained on a visual dictionary – the ImageNet project – with a database of more than 14 million objects. Each object is described by a mathematical term, or vector, that enables the machine to recognise the shape the next time it is encountered. Those mathematical definitions are linked to the words humans would use to describe the objects.
“I was amazed that even with the small amount of training data that we were able to do so well,” said Oriol Vinyals, a Google computer scientist who worked with members of the Google Brain project. “The field is just starting, and we will see a lot of increases.”
In the near term, computer vision systems that can discern the story in a picture will enable people to search photo or video archives and find highly specific images. Eventually, these advances will lead to robotic systems able to navigate unknown situations. Driverless cars would also be made safer. However, it also raises the prospect of even greater levels of government surveillance.

 frisbee 
"A group of young people playing a game of Frisbee."
 

 frisbee 
"A person riding a motorcycle on a dirt road."
 

 frisbee 
"A pizza sitting on top of a pan on top of a stove."
 

Comments

Popular posts from this blog

Use Cortana to define words for you in Windows 10

Microsoft’s Cortana has many uses including sending emails or checking the weather. One of the best uses though is a simple look-up feature for words and their definitions. Combined with the Hey Cortana voice recognition using Cortana to tell you quickly what a word means is a great hands-free tip. There are multiple ways to get Cortana to define a word. If you have Hey Cortana enabled you can simply blurt out your request: “Hey Cortana what is the meaning/definition of  inchoate ?” “Hey Cortana define  ubiquitously “ You can, of course, also just type in your request e.g “ define beatitude ” although this admittedly takes some of the speed (and fun) out of using the personal digital assistant. For many words, Cortana displays the definition within a card, and OxfordDictionaries or EncartaDictionaries powers it. Sometimes, if Cortana misunderstands you or does not have the definition, the assistant opens up a web page after performing a web search for...

Intel announces the first 14 nanometre processor

At the Computex conference in Taipei, chipmaker Intel has revealed a fanless mobile PC reference design using the first of its next-generation 14nm "Broadwell" processors. The 2 in 1 pictured here is a 12.5" screen that is just 7.2 mm thick with keyboard detached and weighs 670 grams.  The Surface Pro 3  – for comparison – is 9.1 mm thick and weighs 800 grams. It includes a media dock that provides additional cooling for a burst of performance. The next-generation chip is purpose-built for 2 in 1s and will hit the market later in  2014 . Called the Intel Core M, it will be the most energy-efficient Intel Core processor in the company's history with power usage cut by up to 45 percent, resulting in 60 percent less heat. The majority of designs based on this new chip are expected to be fanless, with up to  32 hours of battery life,  offering both a lightning-fast tablet and razor-thin laptop. Intel is also delivering innovation and performance for the ...

How ad-free subscriptions could solve Facebook

At the core of Facebook’s “well-being” problem is that its business is directly coupled with total time spent on its apps. The more hours you pass on the social network, the more ads you see and click, the more money it earns. That puts its plan to make using Facebook healthier at odds with its finances, restricting how far it’s willing to go to protect us from the harms of over use. The advertising-supported model comes with some big benefits, though. Facebook CEO Mark Zuckerberg has repeatedly said that “We will always keep Facebook a free service for everyone.” Ads lets Facebook remain free for those who don’t want to pay, and more importantly, for those around the world who couldn’t afford to. Ads pay for Facebook to keep the lights on, research and develop new technologies, and profit handsomely in a way that attracts top talent and further investment. More affluent users with more buying power in markets like the US, UK, and Canada command higher ad prices, effectively...

New wearable tracker can transmit vital signs from a soft, tiny package

Body sensors have long been bulky, hard to wear, and obtrusive. Now they can be as thin as a Band-Aid and about as big as a coin. The new sensors, created by Kyung-In Jang, professor of robotics engineering at South Korea’s Daegu Gyeongbuk Institute of Science and Technology, and John A. Rogers, Northwestern University, consists of a silicone case that contains “50 components connected by a network of 250 tiny wire coils.” The silicone conforms to the body and transmits data on “movement and respiration, as well as electrical activity in the heart, muscles, eyes and brain.” This tiny package replaces many bulky sensor systems and because the wires are suspended in the silicone you are able to create a denser electronic. From the release: Unlike flat sensors, the tiny wires coils in this device are three-dimensional, which maximizes flexibility. The coils can stretch and contract like a spring without breaking. The coils and sensor components are also configured in an unusual spi...