VOOZH
about
URL: https://huggingface.co/monsoon-nlp/activity/community
โฑ monsoon-nlp (Nick Doiron) โ Community Activity
๐ Nick Doiron's picture
51
69
213
Nick Doiron
monsoon-nlp
๐ Image
eramax's profile picture
๐ Image
zhaoxu98's profile picture
๐ Image
AISecHub's profile picture
ยท
https://mapmeld.com/plant-based-llms/
mapmeld
mapmeld.bsky.social
AI & ML interests
biology and multilingual models
Recent Activity
liked
a model
22 days ago
orgava/dna-bacteria-jepa
reacted
to
mmhamdy
's
post
with ๐
22 days ago
Human brains don't recreate every pixel to understand the world! Most current models in genomics, proteomics, and single-cell transcriptomics rely on generative objectives like masked language modeling or next token prediction. While effective, these architectures waste significant capacity reconstructing raw, noisy sequence details that may not carry functional biological meaning. But a promising, more efficient alternative is emerging: Joint-Embedding Predictive Architecture (JEPA) Originally introduced by Yann LeCun for computer vision, JEPA is a non-generative, self-supervised learning (SSL) framework. Instead of predicting raw inputs, it operates as a world model that predicts abstract semantic embeddings in latent space. Recently, the JEPA framework (and its more efficient LeJEPA variant) has been adapted into the biological sciences to develop performing foundation models and to improve on already existing ones. It's interesting how each adaptation modified and tailored JEPA to suit its specific biological domain, whether by experimenting with different backbones or complementing the objective with other loss terms. For example, JEPA-DNA and ProteinJEPA used JEPA as a continual pre-training framework to enhance existing foundation models without training from scratch, while Cell-JEPA and JEPA-DNA employed a hybrid objective that combines the JEPA loss with a traditional language modeling loss. The article below provides an overview of these implementations, along with others that came out this year. As always, your thoughts and feedback are welcome and highly appreciated! Link to the article is in the first comment ๐
reacted
to
pankajpandey-dev
's
post
with ๐ฅ
29 days ago
๐ฎ๐ณ Qwen3-4B Hindi Instruct v2 โ a Hindi LLM that runs on your own machine Most strong Hindi-capable models are either huge or cloud-only. I wanted one that's small enough to run locally but actually follows instructions in Hindi โ so I fine-tuned Qwen3-4B on 10K Hindi instruction pairs and shipped it with a full GGUF quant ladder. โ Fine-tune (16-bit): huggingface.co/pankajpandey-dev/Qwen3-4B-Hindi-Instruct-v2 โ GGUF (Q4/Q5/Q8): huggingface.co/pankajpandey-dev/Qwen3-4B-Hindi-Instruct-v2-GGUF Runs in Ollama, llama.cpp, and LM Studio. The Q4_K_M is just 2.5 GB โ fits comfortably on a laptop, CPU or GPU. Part of my Hindi LLM Series โ building openly-licensed Indic models for local and edge use. More coming (Gemma next). Feedback welcome ๐ #Hindi #IndicNLP #GGUF #LocalLLM #Qwen
View all activity
Organizations
๐ BigScience Workshop's profile picture
๐ Spaces-explorers's profile picture
๐ BigCode's profile picture
๐ Blog-explorers's profile picture
๐ Scary Snake's profile picture
๐ Hugging Face Discord Community's profile picture
๐ Hugging Face Context Course's profile picture
๐ Image
New activity in
scarysnake/outfitter-advice
6 months ago
[bot] Conversion to Parquet
#1 opened 11 months ago by
๐ Image
parquet-converter
๐ Image
New activity in
scarysnake/sensory-awareness-benchmark
11 months ago
Models stop describing their architecture and limits challenge
#3 opened 11 months ago by
๐ Image
๐ Image
monsoon-nlp
Add new models
#2 opened 11 months ago by
๐ Image
๐ Image
monsoon-nlp
๐ Image
New activity in
bigcode/the-stack-v2
11 months ago
๐ฉ Report: Copyright infringement
2
#34 opened 11 months ago by
deleted
๐ Image
New activity in
scarysnake/sensory-awareness-benchmark
11 months ago
[bot] Conversion to Parquet
#1 opened 11 months ago by
๐ Image
parquet-converter
๐ Image
New activity in
scarysnake/code-refusal-for-abliteration
about 1 year ago
Integrate m/re spurces
1
#1 opened over 1 year ago by
๐ Image
๐ Image
monsoon-nlp
๐ Image
New activity in
monsoon-nlp/redcode-hf
over 1 year ago
[bot] Conversion to Parquet
#1 opened over 1 year ago by
๐ Image
parquet-converter
๐ Image
New activity in
monsoon-nlp/gpt-nyc-nontoxic
over 1 year ago
Adding `safetensors` variant of this model
#1 opened over 1 year ago by
๐ Image
SFconvertbot
๐ Image
New activity in
monsoon-nlp/dv-muril
over 1 year ago
Adding `safetensors` variant of this model
#1 opened over 1 year ago by
๐ Image
SFconvertbot
๐ Image
New activity in
monsoon-nlp/dv-labse
over 1 year ago
Adding `safetensors` variant of this model
#1 opened over 1 year ago by
๐ Image
SFconvertbot
๐ Image
New activity in
monsoon-nlp/genetic-counselor-multiple-choice
over 1 year ago
[bot] Conversion to Parquet
#1 opened over 1 year ago by
๐ Image
parquet-converter
๐ Image
New activity in
monsoon-nlp/genetic-counselor-freeform-questions
over 1 year ago
[bot] Conversion to Parquet
#1 opened over 1 year ago by
๐ Image
parquet-converter
๐ Image
New activity in
monsoon-nlp/byt5-basque
over 1 year ago
Adding `safetensors` variant of this model
#1 opened over 1 year ago by
๐ Image
SFconvertbot
๐ Image
New activity in
monsoon-nlp/byt5-dv
over 1 year ago
Adding `safetensors` variant of this model
#1 opened over 1 year ago by
๐ Image
SFconvertbot
๐ Image
New activity in
monsoon-nlp/dv-wave
over 1 year ago
Adding `safetensors` variant of this model
#1 opened over 1 year ago by
๐ Image
SFconvertbot
๐ Image
New activity in
monsoon-nlp/byt5-base-dv
over 1 year ago
Adding `safetensors` variant of this model
#1 opened over 1 year ago by
๐ Image
SFconvertbot
๐ Image
commented
a paper
over 1 year ago
Paper
โข
2501.00663
โข
Published
Dec 31, 2024
โข
31
โข
3
๐ Image
New activity in
monsoon-nlp/bangla-electra
over 1 year ago
Adding `safetensors` variant of this model
#1 opened over 1 year ago by
๐ Image
SFconvertbot
๐ Image
commented
2 papers
over 1 year ago
Paper
โข
2411.07781
โข
Published
Nov 12, 2024
โข
1
โข
1
Paper
โข
2410.20771
โข
Published
Oct 28, 2024
โข
3
โข
1