Cutting a 220,000-token document down to 4,000 with K-means
Part 2 on summarizing a 200,000-word handbook: validating K-means clusters with UMAP and silhouette scores, picking the chunk nearest each centroid, and where the final reduce step still drops topics.