Instructions to use PolyU-AngelChanLab/stanza-yue-hk-electra-large-lora-r8 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Stanza
How to use PolyU-AngelChanLab/stanza-yue-hk-electra-large-lora-r8 with Stanza:
import stanza stanza.download("yue-hk-electra-large-lora-r8") nlp = stanza.Pipeline("yue-hk-electra-large-lora-r8") - Notebooks
- Google Colab
- Kaggle
Cantonese Stanza ELECTRA dependency parser (LoRA r=8)
This is a Cantonese-adapted Transformer dependency parser. Our PhD student Sebastian Chuqiao SONG started from Stanza 1.14.0's Mandarin Chinese zh-hans/gsdsimp_electra-large dependency checkpoint and performed the adaptation.
The ELECTRA base Transformer was adapted with LoRA (r=8, alpha=16, dropout 0.1), while the Stanza parsing layers were fully fine-tuned. The base Transformer weights were frozen. Mandarin POS and lemma processors were also frozen, and their predictions were cached so that the upstream input condition stayed identical before and after parser adaptation.
Intended use
The checkpoint is intended for research on Cantonese Universal Dependencies parsing under this fixed input protocol:
- gold sentence boundaries and gold tokenization;
- frozen Mandarin Stanza predicted POS and lemma features;
- predicted
HEADandDEPREL; - official CoNLL-2018 LAS as the primary metric.
It is not an end-to-end raw-text Cantonese pipeline. It requires the matching Stanza resources and the pinned ELECTRA revision listed below.
Why adaptation was needed
The available Stanza checkpoints evaluated in this experiment were trained for Mandarin Chinese, not Cantonese. On the fixed Cantonese test split, the best unadapted model (zh-hans/gsdsimp_electra-large) reached 41.32% CoNLL-2018 LAS. The adapted checkpoint reached 75.66% LAS.
LoRA was chosen as a parameter-efficient way to adapt the large Transformer with a small Cantonese training set (803 sentences; 11,001 integer-ID word nodes). This is a practical motivation, not evidence that LoRA is intrinsically superior to full fine-tuning. The parsing layers were trained in full, so the resulting system should not be described as changing only low-rank parameters.
Data
- Source: UD Cantonese-HK, release
r2.18 - Source commit:
fcc7dd5b5eb97004441c4aa20704055ddde7a748 - License: CC BY-SA 4.0
- Original treebank: 1,004 sentences and 13,918 integer-ID word nodes
- Custom grouped random split, seed 42: 803 train / 101 dev / 100 test sentences
UD Cantonese-HK officially supplies a test file only. The 803/101/100 split used here is a custom within-treebank experimental split, not an official UD split and not a cross-domain or child-language generalization test. Exact and concatenated FORM duplicate keys were grouped to prevent duplicate text from crossing splits.
No child SRep data were used to train this parser.
Training configuration
| Item | Value |
|---|---|
| Stanza | 1.14.0 |
| Base Stanza dependency package | zh-hans/gsdsimp_electra-large |
| ELECTRA base | hfl/chinese-electra-180g-large-discriminator |
| ELECTRA revision | d017e219578df8e4885484edbc8969dbdea9cbe0 |
| LoRA | r=8, alpha=16, dropout 0.1 |
| LoRA targets | query, value, output.dense, intermediate.dense |
| Parser learning rate | 1e-3 |
| LoRA learning rate | 2e-5 |
| Selection | best dev CoNLL-2018 LAS, evaluated every 100 steps |
| Best step | 1,200 |
If the Cantonese training data contained a dependency relation absent from the Mandarin parser vocabulary, only the new output units were initialized; all old relation slices and other parser weights were preserved at initialization.
Results
| Model | Dev CoNLL-2018 LAS | Test CoNLL-2018 LAS | Test strict full-DEPREL LAS |
|---|---|---|---|
| Adapted checkpoint | 74.82 | 75.66 | 74.78 |
Licensing
This repository is released under CC BY-SA 4.0 to preserve the share-alike terms of the UD Cantonese-HK and UD Chinese-GSDSimp training data. Stanza software is Apache-2.0; Stanza language packs are separately described by Stanford as ODC Attribution 1.0 to the extent Stanford holds rights in them; the ELECTRA base model is Apache-2.0. Users remain responsible for complying with all upstream terms.
References
- Peng Qi, Yuhao Zhang, Yuhui Zhang, Jason Bolton, and Christopher D. Manning. 2020. Stanza: A Python Natural Language Processing Toolkit for Many Human Languages. ACL System Demonstrations.
- Edward J. Hu et al. 2022. LoRA: Low-Rank Adaptation of Large Language Models. ICLR.
- Kevin Clark et al. 2020. ELECTRA: Pre-training Text Encoders as Discriminators Rather Than Generators. ICLR.
- Universal Dependencies Cantonese-HK r2.18.
Citation
Until a paper or archival model release is available, cite the upstream Stanza, LoRA, ELECTRA, and UD Cantonese-HK resources above, and include the Hugging Face repository URL plus its immutable commit hash in the methods or data/code availability statement.
- Downloads last month
- -
Dataset used to train PolyU-AngelChanLab/stanza-yue-hk-electra-large-lora-r8
Evaluation results
- CoNLL-2018 LAS on UD Cantonese-HK r2.18, custom grouped splitself-reported75.662