Profile
Back to NewsBack
GitHub Trending 5 min
Reader Mode
ailia-ai/ailia-models: The collection of pre-trained, state-of-the-art AI models for ailia SDK

ailia-ai/ailia-models: The collection of pre-trained, state-of-the-art AI models for ailia SDK

14 hours ago

GitHub stars</a> PyPI</a> !Python !Platform

The collection of pre-trained, state-of-the-art AI models: 419 models covering object detection, speech recognition, image generation, LLMs and more — all runnable from the same simple CLI.

Tutorial · チュートリアル · Google Colaboratory · Documentation · deepwiki · Update history

Every model works the same way: no arguments needed, weights download automatically.

pip3 install ailia
git clone https://github.com/ailia-ai/ailia-models
cd ailia-models
pip3 install -r requirements.txt
cd object_detection/yolox
python3 yolox.py

Models

419 models are available. Use 🔍 Search models to find a model by name.

| | Category | Model list | |:---|:---|:---| | | Action recognition | va-cnn, st-gcn, mars, ax_action_recognition, driver-action-recognition-adas, action_clip | | | Anomaly detection | mahalanobisad, spade-pytorch, padim, patchcore, glass | | | Audio language model | qwen_audio | | | Audio processing | Audio classification: crnn_audio_classification, audioset_tagging_cnn, transformer-cnn-emotion-recognition, microsoft clap, clap
Music enhancement: hifigan, deep music enhancer
Music generation: pytorch_wavenet
Noise reduction: rnnoise, voicefilter, unet_source_separation, demucs, dtln, voicesplit, audiosep
Phoneme alignment: narabas
Pitch detection: crepe
Speaker diarization: pyannote-audio, auto_speech, wespeaker
Speech to text: deepspeech2, whisper, reazon_speech, distil-whisper, sensevoice, reazon_speech2, kotoba-whisper, lite-whisper
Text to speech: pytorch-dc-tts, tacotron2, vall-e-x, Bert-VITS2, gpt-sovits, gpt-sovits-v2, cosyvoice2, gpt-sovits-v3, gpt-sovits-v2-pro, qwen3-tts
Voice activity detection: silero-vad
Voice conversion: rvc | | | Autonomous driving | bevformer, segformer, uniad | | | Background removal | deep-image-matting, indexnet, U-2-Net, u2net-portrait-matting, u2net-human-seg, cascade_psp, rembg, gfm, modnet, background_matting_v2, dis_seg | | | Crowd counting | crowdcount-cascaded-mtl, c-3-framework | | | Deep fashion | fashionai-key-points-detection, person-attributes-recognition-crossroad, clothing-detection, mmfashion, mmfashion_tryon, mmfashion_retrieval | | | Depth estimation | fcrn-depthprediction, monodepth2, fast-depth, midas, hitnet, lap-depth, mobilestereonet, crestereo, zoe_depth, depth_anything, depth_anything_v2, depth_pro, depth_anything_v3 | | | Diffusion | Text to image: latent-diffusion-txt2img, stable-diffusion-txt2img, anything_v3, control_net, sdxl, latent-consistency-models, sd-turbo, sdxl-turbo, depth_anything_controlnet, latentsync
Text to audio: riffusion
Others: latent-diffusion-inpainting, latent-diffusion-superresolution, DA-CLIP, marigold | | | Face detection | mtcnn, yolov1-face, face-detection-adas, retinaface, blazeface, yolov3-face, face-mask-detection, dbface, anime-face-detector | | | Face identification | facenet_pytorch, insightface, vggface2, arcface, cosface | | | Face recognition | Age gender estimation: face_classification, age-gender-recognition-retail, mivolo, ailia_age_gender
Emotion recognition: ferplus, hsemotion
Gaze estimation: gazeml, mediapipe_iris, gazelle, ax_gaze_estimation
Head pose estimation: hopenet, 6d_repnet, L2CS_Net, 6d_repnet_360
Keypoint detection: face_alignment, prnet, facemesh, facial_feature, 3ddfa, facemesh_v2
Others: face-anti-spoofing, ax_facial_features | | | Face restoration | gfpgan, codeformer | | | Face swapping | deepfacelive, sber-swap, facefusion | | | Feature extraction | dinov3 | | | Frame interpolation | cain, rife, flavr, film | | | Generative adversarial networks | pytorch-gan, lipgan, council-gan, sam, encoder4editing, restyle-encoder, SadTalker, live_portrait | | | Hand detection | hand_detection_pytorch, yolov3-hand, blazepalm | | | Hand recognition | hand3d, v2v-posenet, minimal-hand, blazehand, hands_segmentation_pytorch | | | Image captioning | illustration2vec, image_captioning_pytorch, blip2 | | | Image classification | CNN: alexnet, vgg16, googlenet, resnet18, resnet50, inceptionv3, inceptionv4, wide_resnet50, mobilenetv2, mobilenetv3, efficientnet, efficientnetv2, imagenet21k, mlp_mixer, volo, convnext, mobileone
Transformer: vit, clip, swin-transformer, japanese-clip, japanese-stable-clip-vit-l-16, siglip-multilingual, clip-japanese-base, siglip2
Specific task: weather-prediction-from-image, partialconv | | | Image inpainting | inpainting-with-partial-conv, deepfillv2, inpainting_gmcnn, 3d-photo-inpainting, lama | | | Image manipulation | colorization, cnngeometric_pytorch, style2paints, deblur_gan, pytorch-superpoint, noise2noise, dfe, illnet, dewarpnet, deep_white_balance, u2net_portrait, invertible_denoising_network, dfm, fbcnn, dehamer, lightglue, docshadow | | | Image quality assessment | aesthetic-predictor | | | Image restoration | nafnet | | | Image segmentation | pytorch-fcn, pytorch-enet, tusimple-DUC, pytorch-unet, deeplabv3, pspnet-hair-segmentation, swiftnet, hrnet_segmentation, hair_segmentation, paddleseg, human_part_segmentation, semantic-segmentation-mobilenet-v3, suim, yet-another-anime-segmenter, dense_prediction_transformers, group_vit, pp_liteseg, anime-segmentation, yolov8-seg, segment-anything, grounded_sam, fast_sam, mobile_sam, edge_sam, segment-anything-2, yolov11-seg, segment-anything-3.1 | | | Landmark classification | places365, landmarks_classifier_asia | | | Line segment detection | dexined, mlsd | | | Low light image enhancement | agllnet, drbn_skf | | | Natural language processing | Bert: bert, bert_maskedlm, bert_question_answering
Embedding: sentence_transformers_japanese, multilingual-e5, glucose, qwen3-embedding, ruri-v3, embeddinggemma
Error corrector: bert_insert_punctuation, bertjsc, t5_whisper_medical
Grapheme to phoneme: g2p_en, g2pw, soundchoice-g2p
Named entity recognition: bert_ner, t5_base_japanese_ner, bert_ner_japanese
Reranker: cross_encoder_mmarco, japanese-reranker-cross-encoder, ruri-v3-reranker
Sentence generation: gpt2, rinna_gpt2
Sentiment analysis: bert_sentiment_analysis, bert_tweets_sentiment
Summarize: bert_sum_ext, presumm, t5_base_japanese_title_generation, t5_base_summarization
Translation: fugumt-en-ja, fugumt-ja-en
Zero shot classification: bert_zero_shot_classification, multilingual-minilmv2 | | | Network intrusion detection | bert-network-packet-flow-header-payload, falcon-adapter-network-packet | | | Neural rendering | nerf, TripoSR | | | NSFW detector | clip-based-nsfw-detector | | | Object detection | CNN: yolov1-tiny, yolov2, yolov2-tiny, maskrcnn, yolov3, yolov3-tiny, mobilenet_ssd, m2det, centernet, yolact, efficientdet, pedestrian_detection, crowd_det, yolov4, yolov4-tiny, yolov5, poly_yolo, nanodet, yolor, yolox, picodet, yolox-ti-lite, yolov7, fastest-det, yolov, yolov6, damo_yolo, yolov8, yolox_body_head_hand_face, yolov9, yolov10, yolov11, yolov12
Transformer: detr, glip, dab-detr, detic, groundingdino, rt-detr-v2
Specific target: traffic-sign-detection, sku110k-densedet, footandball, qrcode_wechatqrcode, mobile_object_localizer, layout_parsing | | | Object detection 3d | 3d_bbox, d4lcn, egonet, mediapipe_objectron, 3d-object-detection.pytorch, did_m3d | | | Object tracking | deepsort, person_reid_baseline_pytorch, abd_net, deepsort_vehicle, qd-3dt, centroids-reid, siam-mot, bytetrack, strong_sort, samurai | | | Optical flow estimation | raft, cotracker3 | | | Point segmentation | pointnet_pytorch | | | Pose estimation | openpose, posenet, pose_resnet, lightweight-human-pose-estimation, animalpose, efficientpose, blazepose, mediapipe_holistic, movenet, ap-10k, e2pose | | | Pose estimation 3d | pose-hg-3d, 3d-pose-baseline, lightweight-human-pose-estimation-3d, 3dmppe_posenet, gast, blazepose-fullbody, mediapipe_pose_world_landmarks | | | Road detection | road-segmentation-adas, codes-for-lane-detection, ultra-fast-lane-detection, polylanenet, roneld, lstr, yolop, cdnet, hybridnets | | | Rotation prediction | rotnet | | | Style transfer | adain, pix2pixHD, beauty_gan, psgan, animeganv2, EleGANt | | | Super resolution | srresnet, edsr, han, real-esrgan, swinir, rcan-it, Hat, SPAN | | | Text detection | east, pixel_link, craft_pytorch | | | Text recognition | etl, crnn.pytorch, deep-text-recognition-benchmark, easyocr, paddleocr, donut, ndlocr_text_recognition, paddleocr_v3 | | | Time-series forecasting | informer2020, timesfm, moirai, chronos2 | | | Vehicle recognition | vehicle-attributes-recognition-barrier, vehicle-license-plate-detection-barrier | | | Vision language model | llava, florence2, mobilevlm, llava-jp, qwen2_vl, qwen2.5_vl, qwen3_vl | | | Commercial model | acculus-pose |

About ailia SDK

ailia SDK is a cross-platform, high-speed inference SDK for AI. It supports Windows, Mac, Linux, iOS, Android, Jetson, and Raspberry Pi with GPU acceleration via Vulkan and Metal. Bindings are available for C++, Python, Unity (C#), Kotlin, Rust, and Flutter.

Other platforms

Prototype with ailia MODELS (Python), then deploy to production.

Contact

Chat with me