Cantonese Stanza ELECTRA dependency parser (LoRA r=8)

This is a Cantonese-adapted Transformer dependency parser. Our PhD student Sebastian Chuqiao SONG started from Stanza 1.14.0's Mandarin Chinese zh-hans/gsdsimp_electra-large dependency checkpoint and performed the adaptation.

The ELECTRA base Transformer was adapted with LoRA (r=8, alpha=16, dropout 0.1), while the Stanza parsing layers were fully fine-tuned. The base Transformer weights were frozen. Mandarin POS and lemma processors were also frozen, and their predictions were cached so that the upstream input condition stayed identical before and after parser adaptation.

Intended use

The checkpoint is intended for research on Cantonese Universal Dependencies parsing under this fixed input protocol:

  • gold sentence boundaries and gold tokenization;
  • frozen Mandarin Stanza predicted POS and lemma features;
  • predicted HEAD and DEPREL;
  • official CoNLL-2018 LAS as the primary metric.

It is not an end-to-end raw-text Cantonese pipeline. It requires the matching Stanza resources and the pinned ELECTRA revision listed below.

Why adaptation was needed

The available Stanza checkpoints evaluated in this experiment were trained for Mandarin Chinese, not Cantonese. On the fixed Cantonese test split, the best unadapted model (zh-hans/gsdsimp_electra-large) reached 41.32% CoNLL-2018 LAS. The adapted checkpoint reached 75.66% LAS.

LoRA was chosen as a parameter-efficient way to adapt the large Transformer with a small Cantonese training set (803 sentences; 11,001 integer-ID word nodes). This is a practical motivation, not evidence that LoRA is intrinsically superior to full fine-tuning. The parsing layers were trained in full, so the resulting system should not be described as changing only low-rank parameters.

Data

  • Source: UD Cantonese-HK, release r2.18
  • Source commit: fcc7dd5b5eb97004441c4aa20704055ddde7a748
  • License: CC BY-SA 4.0
  • Original treebank: 1,004 sentences and 13,918 integer-ID word nodes
  • Custom grouped random split, seed 42: 803 train / 101 dev / 100 test sentences

UD Cantonese-HK officially supplies a test file only. The 803/101/100 split used here is a custom within-treebank experimental split, not an official UD split and not a cross-domain or child-language generalization test. Exact and concatenated FORM duplicate keys were grouped to prevent duplicate text from crossing splits.

No child SRep data were used to train this parser.

Training configuration

Item Value
Stanza 1.14.0
Base Stanza dependency package zh-hans/gsdsimp_electra-large
ELECTRA base hfl/chinese-electra-180g-large-discriminator
ELECTRA revision d017e219578df8e4885484edbc8969dbdea9cbe0
LoRA r=8, alpha=16, dropout 0.1
LoRA targets query, value, output.dense, intermediate.dense
Parser learning rate 1e-3
LoRA learning rate 2e-5
Selection best dev CoNLL-2018 LAS, evaluated every 100 steps
Best step 1,200

If the Cantonese training data contained a dependency relation absent from the Mandarin parser vocabulary, only the new output units were initialized; all old relation slices and other parser weights were preserved at initialization.

Results

Model Dev CoNLL-2018 LAS Test CoNLL-2018 LAS Test strict full-DEPREL LAS
Adapted checkpoint 74.82 75.66 74.78

Licensing

This repository is released under CC BY-SA 4.0 to preserve the share-alike terms of the UD Cantonese-HK and UD Chinese-GSDSimp training data. Stanza software is Apache-2.0; Stanza language packs are separately described by Stanford as ODC Attribution 1.0 to the extent Stanford holds rights in them; the ELECTRA base model is Apache-2.0. Users remain responsible for complying with all upstream terms.

References

Citation

Until a paper or archival model release is available, cite the upstream Stanza, LoRA, ELECTRA, and UD Cantonese-HK resources above, and include the Hugging Face repository URL plus its immutable commit hash in the methods or data/code availability statement.

Downloads last month
-
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Dataset used to train PolyU-AngelChanLab/stanza-yue-hk-electra-large-lora-r8

Evaluation results

  • CoNLL-2018 LAS on UD Cantonese-HK r2.18, custom grouped split
    self-reported
    75.662