🔄 In a Training Loop
Minh-Thien Nguyen
minhnguyent546
AI & ML interests
Research interests: Embeddings Models, Image-Text retrieval for Vietnamese, Optimal transport, RAG, and Image classification. Some pet projects on distributed training, training model on TPU, RAG for complex document multiple-choice QA.
Recent Activity
updated a model 4 days ago
minhnguyent546/var2026-wheelhouse published a model 4 days ago
minhnguyent546/var2026-wheelhouse liked a dataset 7 days ago
HuggingFaceCode/stack-v3-trainOrganizations
CoTu @ EXACT-2026
cotu-legal-retriever
cotu-legal-retriever is a family of models optimized for Vietnamese legal retrieval tasks.
-
minhnguyent546/cotu-legal-retriever-Octen-Embedding-4B-stage1
Sentence Similarity • 4B • Updated • 93 -
minhnguyent546/cotu-legal-retriever-Qwen3-Embedding-4B-stage1
Sentence Similarity • 4B • Updated • 86 -
minhnguyent546/cotu-legal-retriever-Qwen3-Embedding-8B-stage1
Sentence Similarity • 8B • Updated • 85 -
minhnguyent546/KaLM-Embedding-Gemma3-12B-2511-tokenizer-for-transformers-v5
Updated
[dataset] image-text datasets
[dataset] text-generation
ViCLIP-OT
ViCLIP-OT: The First Foundation Vision-Language Model for Vietnamese Image–Text Retrieval with Optimal Transport
-
minhnguyent546/ViCLIP-OT
Feature Extraction • 0.2B • Updated • 90 • 3 -
minhnguyent546/ViSigLIP-OT
Feature Extraction • 0.2B • Updated • 74 • 1 -
minhnguyent546/ViCLIP-OT-checkpoints
Feature Extraction • Updated -
ViCLIP-OT: The First Foundation Vision-Language Model for Vietnamese Image-Text Retrieval with Optimal Transport
Paper • 2602.22678 • Published
var-2026
ViREx-Bench
CoTu @ EXACT-2026
e2026
cotu-legal-retriever
cotu-legal-retriever is a family of models optimized for Vietnamese legal retrieval tasks.
-
minhnguyent546/cotu-legal-retriever-Octen-Embedding-4B-stage1
Sentence Similarity • 4B • Updated • 93 -
minhnguyent546/cotu-legal-retriever-Qwen3-Embedding-4B-stage1
Sentence Similarity • 4B • Updated • 86 -
minhnguyent546/cotu-legal-retriever-Qwen3-Embedding-8B-stage1
Sentence Similarity • 8B • Updated • 85 -
minhnguyent546/KaLM-Embedding-Gemma3-12B-2511-tokenizer-for-transformers-v5
Updated
[model] Machine Translation Models
[dataset] image-text datasets
[dataset] embeddings-and-retrieval-learning
Datasets for training embeddings models (and fine-tuning for retrieval tasks)
[dataset] text-generation
[model] embeddings
ViCLIP-OT
ViCLIP-OT: The First Foundation Vision-Language Model for Vietnamese Image–Text Retrieval with Optimal Transport
-
minhnguyent546/ViCLIP-OT
Feature Extraction • 0.2B • Updated • 90 • 3 -
minhnguyent546/ViSigLIP-OT
Feature Extraction • 0.2B • Updated • 74 • 1 -
minhnguyent546/ViCLIP-OT-checkpoints
Feature Extraction • Updated -
ViCLIP-OT: The First Foundation Vision-Language Model for Vietnamese Image-Text Retrieval with Optimal Transport
Paper • 2602.22678 • Published
Med-Alpaca