dotlah! dotlah!
  • Cities
  • Technology
  • Business
  • Politics
  • Society
  • Science
  • About
Social Links
  • zedreviews.com
  • citi.io
  • aster.cloud
  • liwaiwai.com
  • guzz.co.uk
  • atinatin.com
0 Likes
0 Followers
0 Subscribers
dotlah!
  • Cities
  • Technology
  • Business
  • Politics
  • Society
  • Science
  • About
  • Machine Learning
  • Research
  • Science
  • Technology

Computer Vision System Marries Image Recognition And Generation

  • July 3, 2023
MIT MAGE
A unified vision system known as MAsked Generative Encoder (MAGE), developed by researchers at MIT and Google, could be useful for many things, like finding and classifying objects in an image, learning from just a few examples, generating images with specific conditions such as text or class, editing existing images, and more. Image: Alex Shipps/MIT CSAIL via Midjourney
Total
0
Shares
0
0
0

MAGE merges the two key tasks of image generation and recognition, typically trained separately, into a single system.

Rachel Gordon | MIT CSAIL

MIT MAGE
A unified vision system known as MAsked Generative Encoder (MAGE), developed by researchers at MIT and Google, could be useful for many things, like finding and classifying objects in an image, learning from just a few examples, generating images with specific conditions such as text or class, editing existing images, and more. Image: Alex Shipps/MIT CSAIL via Midjourney

Computers possess two remarkable capabilities with respect to images: They can both identify them and generate them anew. Historically, these functions have stood separate, akin to the disparate acts of a chef who is good at creating dishes (generation), and a connoisseur who is good at tasting dishes (recognition).

Yet, one can’t help but wonder: What would it take to orchestrate a harmonious union between these two distinctive capacities? Both chef and connoisseur share a common understanding in the taste of the food. Similarly, a unified vision system requires a deep understanding of the visual world.

Now, researchers in MIT’s Computer Science and Artificial Intelligence Laboratory (CSAIL) have trained a system to infer the missing parts of an image, a task that requires deep comprehension of the image’s content. In successfully filling in the blanks, the system, known as the Masked Generative Encoder (MAGE), achieves two goals at the same time: accurately identifying images and creating new ones with striking resemblance to reality. 

This dual-purpose system enables myriad potential applications, like object identification and classification within images, swift learning from minimal examples, the creation of images under specific conditions like text or class, and enhancing existing images.

Unlike other techniques, MAGE doesn’t work with raw pixels. Instead, it converts images into what’s called “semantic tokens,” which are compact, yet abstracted, versions of an image section. Think of these tokens as mini jigsaw puzzle pieces, each representing a 16×16 patch of the original image. Just as words form sentences, these tokens create an abstracted version of an image that can be used for complex processing tasks, while preserving the information in the original image. Such a tokenization step can be trained within a self-supervised framework, allowing it to pre-train on large image datasets without labels. 

Now, the magic begins when MAGE uses “masked token modeling.” It randomly hides some of these tokens, creating an incomplete puzzle, and then trains a neural network to fill in the gaps. This way, it learns to both understand the patterns in an image (image recognition) and generate new ones (image generation).

“One remarkable part of MAGE is its variable masking strategy during pre-training, allowing it to train for either task, image generation or recognition, within the same system,” says Tianhong Li, a PhD student in electrical engineering and computer science at MIT, a CSAIL affiliate, and the lead author on a paper about the research. “MAGE’s ability to work in the ‘token space’ rather than ‘pixel space’ results in clear, detailed, and high-quality image generation, as well as semantically rich image representations. This could hopefully pave the way for advanced and integrated computer vision models.” 

Apart from its ability to generate realistic images from scratch, MAGE also allows for conditional image generation. Users can specify certain criteria for the images they want MAGE to generate, and the tool will cook up the appropriate image. It’s also capable of image editing tasks, such as removing elements from an image while maintaining a realistic appearance.

Recognition tasks are another strong suit for MAGE. With its ability to pre-train on large unlabeled datasets, it can classify images using only the learned representations. Moreover, it excels at few-shot learning, achieving impressive results on large image datasets like ImageNet with only a handful of labeled examples.

The validation of MAGE’s performance has been impressive. On one hand, it set new records in generating new images, outperforming previous models with a significant improvement. On the other hand, MAGE topped in recognition tasks, achieving an 80.9 percent accuracy in linear probing and a 71.9 percent 10-shot accuracy on ImageNet (this means it correctly identified images in 71.9 percent of cases where it had only 10 labeled examples from each class).

Despite its strengths, the research team acknowledges that MAGE is a work in progress. The process of converting images into tokens inevitably leads to some loss of information. They are keen to explore ways to compress images without losing important details in future work. The team also intends to test MAGE on larger datasets. Future exploration might include training MAGE on larger unlabeled datasets, potentially leading to even better performance. 

“It has been a long dream to achieve image generation and image recognition in one single system. MAGE is a groundbreaking research which successfully harnesses the synergy of these two tasks and achieves the state-of-the-art of them in one single system,” says Huisheng Wang, senior staff software engineer of humans and interactions in the Research and Machine Intelligence division at Google, who was not involved in the work. “This innovative system has wide-ranging applications, and has the potential to inspire many future works in the field of computer vision.” 

Li wrote the paper along with Dina Katabi, the Thuan and Nicole Pham Professor in the MIT Department of Electrical Engineering and Computer Science and a CSAIL principal investigator; Huiwen Chang, a senior research scientist at Google; Shlok Kumar Mishra, a University of Maryland PhD student and Google Research intern; Han Zhang, a senior research scientist at Google; and Dilip Krishnan, a staff research scientist at Google. Computational resources were provided by Google Cloud Platform and the MIT-IBM Watson AI Lab. The team’s research was presented at the 2023 Conference on Computer Vision and Pattern Recognition.

Reprinted with permission of MIT News (http://news.mit.edu/)

Total
0
Shares
Share
Tweet
Share
Share
Related Topics
  • Google
  • image generation
  • Machine Learning
  • MAGE
  • MAsked Generative Encoder
  • MIT
  • MIT CSAIL
John Francis

Previous Article
USA flag
  • Featured
  • People

Stars, Stripes, And Service. Exploring The Hierarchies And Heroes Of The U.S. Military

  • July 3, 2023
View Post
Next Article
usa-flag-justin-cron-_gtwjIzQLq4-unsplash
  • Features
  • People
  • World Events

Stars, Stripes, And Service. Exploring The Hierarchies And Heroes Of The U.S. Military

  • July 3, 2023
View Post
You May Also Like
View Post
  • Artificial Intelligence
  • Technology

Google Opens Singapore Engineering Center to Build and Export Enterprise Cloud and AI to the World

  • Dean Marc
  • September 15, 2026
View Post
  • Research
  • Technology

MIT researchers tackle the economic realities of fusion power

  • dotlah.com
  • August 10, 2026
zedreviews-valerion
View Post
  • Gears
  • Technology

Father’s Day Outdoors – Build Dad the Ultimate Backyard Watch Party

  • Dean Marc
  • June 21, 2026
zedreviews-fathers-day-50830
View Post
  • Gears
  • Technology

Father’s Day Outdoors, Round Two – Gear for the Action, the Tailgate, and Beating the Heat

  • Dean Marc
  • June 21, 2026
zedreviews-fathers-day-21306
View Post
  • Gears
  • Technology

A Father’s Day Gift Guide for Every Dad – Timepieces and Travel Gear

  • Dean Marc
  • June 21, 2026
View Post
  • Gears
  • Technology

Samsung Art Store Brings Art Basel to Homes Worldwide With New Curated Collection

  • Dean Marc
  • June 15, 2026
View Post
  • Artificial Intelligence
  • Technology

The consequences of relying on AI for accurate news

  • dotlah.com
  • June 10, 2026
pope-leo-xiv-cq5dam-1500.844
View Post
  • Artificial Intelligence
  • Technology

Pope Leo XIV to Publish First Encyclical on Artificial Intelligence and Human Dignity on 25 May

  • dotlah.com
  • May 23, 2026


Trending
  • pandemic-crowd-macau-photo-agency-4yXV0JIK-yo-unsplash 1
    • People
    • World Events
    How many people need to get a COVID-19 vaccine in order to stop the coronavirus?
    • January 17, 2021
  • 2
    • Lah!
    ACRA And SGX RegCo Emphasise Importance Of High-quality Financial Statements Amid COVID-19
    • July 29, 2020
  • 3
    • Lah!
    ​Sea Level Could Rise By More Than 1 Metre By 2100 If Emission Targets Are Not Met, Reveals Survey Of 100 International Experts
    • May 12, 2020
  • 4
    • Cities
    • Society
    Asia Dominates When It Comes to Passport Power In 2020
    • January 14, 2020
  • 5
    • Society
    UOB Group Commercial Banking And Clients Spring Into Action And Raise More Than $1.8 million At UOB’s Annual Lunar New Year Fundraiser
    • January 30, 2020
  • 6
    • Lah!
    Pinterest Opens Its Doors In Singapore
    • July 11, 2019
  • 7
    • Cities
    • Lah!
    MPA Launches Regular Guided Tours To Raffles Lighthouse
    • January 31, 2022
  • 8
    • Science
    • Society
    COVID-19 Treatment Might Already Exist In Old Drugs – We’re Using Pieces Of The Coronavirus Itself To Find Them
    • March 20, 2020
  • 9
    • People
    • World Events
    The Mysterious Disappearance Of The First SARS Virus, And Why We Need A Vaccine For The Current One But Didn’t For The Other
    • May 5, 2020
  • 10
    • Cities
    • People
    How We Took Everything For Granted – And Will We Do It Again?
    • May 21, 2020
  • 11
    • Lah!
    New Changi Experience Studio in Jewel Brings Visitors On a Journey of Fun & Discovery
    • June 10, 2019
  • 12
    • Cities
    • Society
    Keppel Announces $4.2 Million Package To Support National Efforts To Combat COVID-19
    • March 20, 2020
Trending
  • 1
    Chinese tech giant Huawei unveils Mate 90 series smartphones with flagship Tau chips adopting logic folding technology
    • October 3, 2026
  • 2
    Google Opens Singapore Engineering Center to Build and Export Enterprise Cloud and AI to the World
    • September 15, 2026
  • 3
    Apple Flagship Hardware Launch for September 2026
    • September 10, 2026
  • 4
    MIT researchers tackle the economic realities of fusion power
    • August 10, 2026
  • 5
    ASEAN, treaty parties push stronger cooperation as Treaty of Amity and Cooperation in Southeast Asia marks 50 years
    • July 26, 2026
  • 6
    The Fastest AI Fried Chicken In The World
    • June 29, 2026
  • 7
    Zed Approves | How to Stay Cool in Extreme Heat
    • June 29, 2026
  • 8
    Was Venezuela struck by an earthquake ‘doublet’? Here’s what we know so far
    • June 27, 2026
  • 9
    Zed Approves | It’s Prime Day 2026! Time to Upgrade Your World Cup Viewing Setup and Beat the Heat
    • June 25, 2026
  • 10
    Zed Approves | The Best Prime Day PC Deals: Top Gaming Rigs, Workstations, and Everyday Laptops
    • June 25, 2026
Social Links
dotlah! dotlah!
  • Cities
  • Technology
  • Business
  • Politics
  • Society
  • Science
  • About
Connecting Dots Across Asia's Tech and Urban Landscape

Input your search keywords and press Enter.