ai

Analysing 200 hours of podcasts with LLMs

There is so much valuable information in podcasts, but so little time to consume and internalize it. I figured I could use LLMs to help me compress the content and surface only the takeaways I care about.

Last week, I put that to the test. In one hour and for just 8 euros, I analyzed more than 200 hours of my favorite podcast and built a knowledge graph + embeddings layer that can be queried via an LLM.

The barrier to entry for experiments like this is dropping fast. Using Cognee (open source: https://github.com/topoteretes/cognee), six lines of code and a MacBook Air were all that was needed to build a proof of concept.

In a nutshell, Cognee tries to overcome some of the limitations of RAG by creating a “memory layer” that combines embeddings with entities stored in a knowledge graph.

For personal projects, it works remarkably well. Queries like: “how to validate a startup”, “how to scale it”, or “how to hire efficiently” successfully surface the information I want. I can even configure the system to reference the specific podcast episode from which each piece of information is derived.

Of course, productionizing is a completely different story. Being an early adopter of relatively new frameworks in a fast-moving AI landscape comes with real risks. This is when you need to dive into the implementation details, understanding its versatility and trade-offs, and evaluating model costs and performance.

Reinventing the wheel isn’t ideal, but realizing too late that a framework is too opinionated and misaligned with your requirements can mean rewriting everything from scratch.