NeuML/domain-labeler

🤗 Hugging Face 来源text-classificationapache-2.032M 参数128 MBsafetensors✓ 1 个校验和今天更新
一条命令提交

在你的模型文件夹旁边运行它。它会制作种子、将文件与 Hugging Face 比对,然后提交。你只需开始做种,并粘贴你账户中的密钥。它只读取你的文件,绝不修改。如果愿意,可以先阅读脚本。

curl -fsSL https://pirateface.co/package.sh | bash -s -- --repo NeuML/domain-labeler ./model-folder
需要做种者 →

Domain Labeler

This is a Ettin 32M parameter model fined-tuned with the Wikipedia Domain Labels dataset for text classification.

This model classifies text into one of the following classes.

labels = [
    "aerospace", "agronomy", "artistic", "astronomy", "atmospheric_science", "automotive", "beauty",
    "biology", "celebrity", "chemistry", "civil_engineering", "communication_engineering",
    "computer_science_and_technology", "design", "drama_and_film", "economics",
    "electronic_science", "entertainment", "environmental_science", "fashion", "finance",
    "food", "gamble", "game", "geography", "health", "history", "hobby", "hydraulic_engineering", 
    "instrument_science", "journalism_and_media_communication", "landscape_architecture", "law",
    "library", "literature", "materials_science", "mathematics", "mechanical_engineering",
    "medical", "mining_engineering", "movie", "music_and_dance", "news", "nuclear_science",
    "ocean_science", "optical_engineering", "painting", "pet", 
    "petroleum_and_natural_gas_engineering", "philosophy", "photo", "physics", "politics",
    "psychology", "public_administration", "relationship", "religion", "sociology", "sports",
    "statistics", "systems_science", "textile_science", "topicality", "transportation_engineering",
    "travel", "urban_planning", "vulgar_language"
]

Usage (txtai)

This model can be used to classify text into one of the domain labels above with txtai.

from txtai.pipeline import Labels

labels = Labels("NeuML/domain-labeler", dynamic=False)
labels("Text to classify")

# Get only the top label
labels("Text to classify", flatten=True)

Usage (Hugging Face Transformers)

The following code is used to run a transformers text-classification pipeline.

labels = pipeline("text-classification", model="NeuML/domain-labeler")
labels("Text to classify")

Evaluation

The following are the metrics for the test dataset. Note that these labels have significant overlap and the overall accuracy is much higher when generalizing the categories. In other words the "wrong" labels aren't always necessarily wrong (i.e. Medical vs Health, Entertainment vs Celebrity etc)

Accuracy F1 Precision Recall PR-ACU
0.8426 83.97 83.96 84.26 90.033

Training code

The training code used to build this model is here.