top of page

Kimi Delta Attention for Full-Document Legal Classification
I conducted a full-document diagnostic experiment examining whether Kimi Delta Attention (KDA) can be practically integrated with LegalBERT for long legal-document classification outside industrial-scale training environments. The task uses state-court cases involving equal-protection and education-adequacy claims.
 
Corrected architecture
The initial implementation was revised after testing showed that tokenizer overflow handling did not reliably preserve every text window in extremely long LexisNexis-derived case files. The corrected pipeline manually tokenizes each document into overlapping 512-token LegalBERT windows and verifies the expected number of chunks before training.

Full court document -> Overlapping 512-token LegalBERT windows -> Pretrained LegalBERT encodes each local legal-text segment -> One legal-aware vector per segment -> Kimi Delta Attention integrates the ordered segment sequence -> Mask-aware document pooling -> One document-level classification prediction

LegalBERT remains unchanged as the local legal-language encoder. KDA is applied after LegalBERT, across the sequence of chunk representations, rather than replacing LegalBERT’s pretrained internal attention layers. This creates a hierarchical full-document model: LegalBERT captures local legal context, while KDA models relationships among segments across the document.
 
Dataset and evaluation
The experiment used 196 retained state-court cases in three categories:

  • Equal protection

  • Education adequacy

  • Both equal protection and education adequacy

A fourth residual category was excluded because it contained only two observations, which was insufficient for reliable stratified evaluation. The model was evaluated with five-fold stratified cross-validation. Each fold trained a newly initialized model for three epochs, producing 15 total fold-epochs. The complete run required approximately 30 minutes on accessible GPU infrastructure.
The implementation uses frozen pretrained LegalBERT for memory-safe local encoding, while KDA and the document-level classifier are trained on the resulting full-document chunk sequences.
 
Verified document coverage
The revised implementation verifies that long documents are fully represented rather than silently truncated to an initial prefix. Because LegalBERT accepts up to 512 tokens per local input, long documents are divided into overlapping windows with 128 tokens of retained overlap.
For example, documents ranging from approximately 13,000 to 70,000 LegalBERT wordpieces produced approximately 35 to 185 local windows. LegalBERT encodes each window, and KDA receives the resulting ordered sequence of 35–185 chunk representations before the model produces one document-level prediction.
A wordpiece is a LegalBERT token unit, not necessarily a complete English word. Legal citations, names, punctuation, section numbers, and compound legal terms may consist of multiple wordpieces.
 
Interpretation and limitations
The predictions are exploratory and should not be treated as definitive legal-classification accuracy. The underlying category labels are literature-derived and preliminary; they are not adjudicated expert gold-standard annotations.
A disagreement between the model and the assigned category may reflect:

  • A model classification error

  • A case involving mixed or ambiguous constitutional claims

  • A coarse, incomplete, or contestable original source label

The study therefore evaluates the feasibility of full-document legal representation and cross-chunk integration under weak supervision. It does not establish a definitive automated classifier for constitutional doctrine.
The dataset is relatively small and imbalanced. Future work should compare this architecture against a matched LegalBERT-only chunk-pooling baseline, use larger and more carefully annotated legal corpora, and include expert review of disputed classifications.
 
Comparative context
KDA belongs to a broader group of efficient-attention approaches intended to reduce conventional softmax attention’s quadratic growth with sequence length while retaining useful long-range information.
Relevant work in the current long-context landscape includes:

  • Kimi Linear / Kimi Delta Attention, which combines KDA with a hybrid attention architecture

  • Ring-linear and other hybrid linear-plus-softmax approaches for efficient long-context processing

  • Log-linear attention approaches that seek to balance efficiency and expressive capacity

  • Nested Learning / Hope-style memory and multi-level optimization architectures

  • Gated delta-rule and gated-linear-attention approaches that use learned state updates and decay mechanisms

This project does not claim to reproduce proprietary architectures used by commercial systems such as GPT, Gemini, or Claude. Instead, it provides an open and reproducible test of a narrower question: whether pretrained legal-language representations can be integrated across an entire judicial document through a KDA-based document-level sequence layer.
 
Takeaway
The corrected experiment establishes a full-document LegalBERT–KDA workflow:

Manual verified document segmentation -> Pretrained LegalBERT encoding of every segment -> KDA-based cross-segment integration -> Mask-aware document pooling -> One final prediction per case
Its primary contribution is methodological rather than a final performance claim. The experiment demonstrates that KDA can be integrated with LegalBERT to analyze complete long-form legal documents on accessible GPU hardware, while making the limitations of small samples, weak source labels, class imbalance, and frozen-encoder training explicit.

bottom of page