microsoft/Phi-4-multimodal-instruct
microsoft/Phi-4-multimodal-instruct : reconnaissance vocale, licence MIT License, 11,9 Gio de poids, 5,6 milliards de paramètres. Mémoire nécessaire estimée,
La source ne fournit pas de description structurée dans les métadonnées consultées.
Fiche factuelle
- Propriétaire déclaré
- microsoft
- Tâche déclarée
- Reconnaissance vocale
automatic-speech-recognition - Langues déclarées
- Multilingue, Arabe, zh, cs, da, nl, Anglais, fi, Français, de, he, hu, it, ja, ko, no, pl, Portugais, ru, es, sv, th, tr, uk
- Licence déclarée
- MIT License
- Qualification de licence
- Licence ouverte permissive
- Usage commercial
- Possible d’après la licence déclarée, à confirmer dans son texte (vérification du texte intégral requise)
- Accès public direct
- Oui, d’après l’API
- Format SafeTensors
- Déclaré
- Bibliothèque
- transformers
- Paramètres
- 5,6 milliards de paramètres
- Taille des poids connue
- 11,9 Gio
- Fichiers référencés
- 55
- Jeux de données déclarés
- Non renseignés
- Résultats d’évaluation structurés
- 0 (comparables seulement à protocole identique)
- Documentation structurée
- Présente
- Dernière mise à jour
- 2025-12-10T20:18:10.000Z
- Téléchargements
- 195684 (indicateur de popularité uniquement)
- Appréciations
- 1616 (indicateur de popularité uniquement)
- Source
- Hugging Face Hub
- Dépôt de code explicitement lié
- Non renseigné
- Publication explicitement liée
- Consulter la publication
- Interrogé le
- 2026-10-09T01:32:08.269487+00:00
Fichiers déclarés par la source
| Nom | Format | Précision détectée | Taille |
|---|---|---|---|
| .gitattributes | autre | inconnue | 1,6 Kio |
| CODE_OF_CONDUCT.md | autre | inconnue | 444 octets |
| LICENSE | autre | inconnue | 1,1 Kio |
| README.md | autre | inconnue | 63,9 Kio |
| SECURITY.md | autre | inconnue | 2,6 Kio |
| SUPPORT.md | autre | inconnue | 1,2 Kio |
| added_tokens.json | autre | inconnue | 249 octets |
| config.json | autre | inconnue | 4,5 Kio |
| configuration_phi4mm.py | autre | inconnue | 10,8 Kio |
| data_summary_card.md | autre | inconnue | 4,5 Kio |
| examples/what_is_shown_in_this_image.wav | autre | inconnue | 110,2 Kio |
| examples/what_is_the_traffic_sign_in_the_image.wav | autre | inconnue | 724,1 Kio |
| figures/audio_understand.png | autre | inconnue | 41,6 Kio |
| figures/multi_image.png | autre | inconnue | 187,5 Kio |
| figures/speech_qa.png | autre | inconnue | 45,7 Kio |
| figures/speech_recog_by_lang.png | autre | inconnue | 88,6 Kio |
| figures/speech_recognition.png | autre | inconnue | 62,1 Kio |
| figures/speech_summarization.png | autre | inconnue | 40,0 Kio |
| figures/speech_translate.png | autre | inconnue | 46,6 Kio |
| figures/speech_translate_2.png | autre | inconnue | 45,3 Kio |
| figures/vision_radar.png | autre | inconnue | 169,7 Kio |
| generation_config.json | autre | inconnue | 190 octets |
| merges.txt | autre | inconnue | 2,3 Mio |
| model-00001-of-00003.safetensors | safetensors | inconnue | 4,7 Gio |
| model-00002-of-00003.safetensors | safetensors | inconnue | 4,6 Gio |
| model-00003-of-00003.safetensors | safetensors | inconnue | 1,1 Gio |
| model.safetensors.index.json | autre | inconnue | 234,3 Kio |
| modeling_phi4mm.py | autre | inconnue | 113,3 Kio |
| phi_4_mm.tech_report.02252025.pdf | autre | inconnue | 5,0 Mio |
| preprocessor_config.json | autre | inconnue | 482 octets |
| processing_phi4mm.py | autre | inconnue | 32,0 Kio |
| processor_config.json | autre | inconnue | 121 octets |
| sample_finetune_speech.py | autre | inconnue | 16,3 Kio |
| sample_finetune_vision.py | autre | inconnue | 19,2 Kio |
| sample_inference_phi4mm.py | autre | inconnue | 10,3 Kio |
| special_tokens_map.json | autre | inconnue | 473 octets |
| speech-lora/adapter_config.json | autre | inconnue | 465 octets |
| speech-lora/adapter_model.safetensors | safetensors | inconnue | 880,0 Mio |
| speech-lora/added_tokens.json | autre | inconnue | 249 octets |
| speech-lora/special_tokens_map.json | autre | inconnue | 473 octets |
| speech-lora/tokenizer.json | autre | inconnue | 14,8 Mio |
| speech-lora/tokenizer_config.json | autre | inconnue | 3,2 Kio |
| speech-lora/vocab.json | autre | inconnue | 3,7 Mio |
| speech_conformer_encoder.py | autre | inconnue | 107,9 Kio |
| tokenizer.json | autre | inconnue | 14,8 Mio |
| tokenizer_config.json | autre | inconnue | 3,2 Kio |
| vision-lora/adapter_config.json | autre | inconnue | 464 octets |
| vision-lora/adapter_model.safetensors | safetensors | inconnue | 704,0 Mio |
| vision-lora/added_tokens.json | autre | inconnue | 249 octets |
| vision-lora/special_tokens_map.json | autre | inconnue | 473 octets |
| vision-lora/tokenizer.json | autre | inconnue | 14,8 Mio |
| vision-lora/tokenizer_config.json | autre | inconnue | 3,2 Kio |
| vision-lora/vocab.json | autre | inconnue | 3,7 Mio |
| vision_siglip_navit.py | autre | inconnue | 76,4 Kio |
| vocab.json | autre | inconnue | 3,7 Mio |