Textual data has always been the hardest data to analyse at scale. NVivo breaks. Human coding does not scale. Keyword search finds words but misses meaning.
Embeddings change this completely.
We were recently exploring a corpus of 26,000 East African job descriptions: over ten million tokens of text. Not analysing yet. Just getting a feel for what was there, using tabular attributes to filter and visualise the vectors in three dimensions.
One question came up naturally: if I got a PhD today, where would I actually work in this market?
The answer was interesting. Most job descriptions do not require a PhD at all. The ones that do cluster into two places almost exclusively: specialist roles, predominantly in NGOs and consulting, and education: lecturers and supervisors. The Kenyan job market is simply not structured around doctoral qualifications outside those two worlds.
But here is where it gets more interesting.
Look at the education cluster more closely and most lecturer descriptions sit tightly together. Similar language, similar requirements, similar expectations. But there are outliers: a small number of lecturer roles sitting semantically far from the rest.
That distance means something. The job title is the same. The underlying need is not.
Those roles are asking for something meaningfully different from what most teaching positions ask for. That kind of observation is invisible to keyword search. It only appears when you let the geometry of the data speak.