Parts-of-Speech as Emergent Categories in SAE Latent Space
This work investigates how Sparse AutoEncoders (SAEs) capture linguistic structure by using part-of-speech (PoS) categories as a controlled test case. The authors find that PoS distinctions are highly recoverable from SAE activations, though they do not align with one-to-one mappings. Instead, morpho-syntactic information is localized in a distributed and category-dependent form, supported by stable, compact groups of sparse latents rather than atomic grammatical features.
PoS distinctions are highly recoverable from SAE activations, but do not align with one-to-one latent / category mappings.
Categories are supported by compact groups of sparse latents, with substantial variation across tags.
These groups remain stable on held-out data, while also showing overlap between related categories.