Text-to-Speech โข Updated
Mohammed Hamdy
๐ Image
ydnysh's profile picture๐ Image
fractalego's profile picture๐ Image
MN0099's profile picture
ydnysh's profile picture๐ Image
fractalego's profile picture๐ Image
MN0099's profile picture
ยท
AI & ML interests
AI4Sci | NLP | Reinforcement Learning
Recent Activity
upvoted a collection about 20 hours ago
๐ Research & Long-Form Blog Posts upvoted an article about 20 hours ago
Profiling in PyTorch (Part 1): A Beginner's Guide to torch.profiler reacted to theirpost with ๐ 1 day ago
What if you could train a model on just 10 images instead of 60,000 and still get close to the same performance?
Traditional machine learning requires thousands, even millions, of data points to achieve high accuracy. But what if we could "distill" the entire dataset into just a few synthetic samples?
This is what Dataset Distillation offers. Unlike traditional knowledge distillation, we keep the model fixed and distill the knowledge contained in a massive training set into a tiny set of synthetic distilled images.
The goal is to train a model on this ultra-small set and achieve performance that almost matches what the same model would get when trained on the massive original dataset.
For example, training on only 10 distilled MNIST images (this is equivalent to a single image per class) yields 94% accuracy, compared to 99% when training on the full 60,000 images.
Interestingly, these distilled images look significantly different (as you can see in the image below) from natural images because they are optimized for model training rather than for matching the correct data distribution.
But that's not all.
Most importantly, this same method opens the door to a potent form of data poisoning. Because distilled images are specifically optimized for rapid learning, an attacker can create a tiny set of adversarial distilled images to cause a well-trained model to forget or misclassify a specific category.
What I find fascinating about dataset distillation is this: it mimics human-like learning by letting a model grasp a concept from a single example, but it does so using alien synthetic images that mean absolutely nothing to a human eye!
What about you? What are your thoughts on it?
Organizations
๐ BigScience Biomedical Datasets's profile picture
๐ fast.ai community's profile picture
๐ Massive Text Embedding Benchmark's profile picture
๐ Blog-explorers's profile picture
๐ Hugging Face for Computer Vision's profile picture
๐ ASAS AI's profile picture
๐ ZeroGPU Explorers's profile picture
๐ Social Post Explorers's profile picture
๐ Cohere Labs Community's profile picture
๐ M4-ai's profile picture
๐ LLMem's profile picture
๐ Hugging Face Discord Community's profile picture
๐ llmc's profile picture
๐ Dataset Tools's profile picture
๐ open/ acc's profile picture
๐ Data Is Better Together Contributor's profile picture
๐ LiteRT Community (FKA TFLite)'s profile picture
๐ MOTH Lab's profile picture
๐ Bitsandbytes Community's profile picture
๐ Reasoning datasets competition 's profile picture
๐ LeRobot Worldwide Hackathon's profile picture
๐ Hugging Face Context Course's profile picture
๐ Agents-MCP-Hackathon's profile picture
๐ Robotics Course's profile picture
๐ Hugging Science's profile picture
๐ Bioscope's profile picture
๐ MCP-1st-Birthday's profile picture
๐ nanochat students's profile picture
๐ Hugging Face Skills's profile picture
๐ Humanity's Last Hackathon's profile picture
๐ Build Small Hackathon's profile picture
๐ fast.ai community's profile picture
๐ Massive Text Embedding Benchmark's profile picture
๐ Blog-explorers's profile picture
๐ Hugging Face for Computer Vision's profile picture
๐ ASAS AI's profile picture
๐ ZeroGPU Explorers's profile picture
๐ Social Post Explorers's profile picture
๐ Cohere Labs Community's profile picture
๐ M4-ai's profile picture
๐ LLMem's profile picture
๐ Hugging Face Discord Community's profile picture
๐ llmc's profile picture
๐ Dataset Tools's profile picture
๐ open/ acc's profile picture
๐ Data Is Better Together Contributor's profile picture
๐ LiteRT Community (FKA TFLite)'s profile picture
๐ MOTH Lab's profile picture
๐ Bitsandbytes Community's profile picture
๐ Reasoning datasets competition 's profile picture
๐ LeRobot Worldwide Hackathon's profile picture
๐ Hugging Face Context Course's profile picture
๐ Agents-MCP-Hackathon's profile picture
๐ Robotics Course's profile picture
๐ Hugging Science's profile picture
๐ Bioscope's profile picture
๐ MCP-1st-Birthday's profile picture
๐ nanochat students's profile picture
๐ Hugging Face Skills's profile picture
๐ Humanity's Last Hackathon's profile picture
๐ Build Small Hackathon's profile picture
Automatic Speech Recognition โข Updated โข 1
Audio Classification โข Updated
Reinforcement Learning โข Updated
Reinforcement Learning โข Updated
Reinforcement Learning โข Updated
Reinforcement Learning โข Updated
Reinforcement Learning โข Updated
Reinforcement Learning โข Updated
Reinforcement Learning โข Updated โข 22
Reinforcement Learning โข Updated โข 2
Reinforcement Learning โข Updated
Reinforcement Learning โข Updated โข 3
Reinforcement Learning โข Updated
Reinforcement Learning โข Updated
Reinforcement Learning โข Updated
Reinforcement Learning โข Updated
