๐๏ธ Generating a Podcast from the Hacker News Front Page
I ain't reading all that. I'm happy for u tho. Or sorry that happened.

Hacker News, most mornings, is thirty stories and a few thousand comments, and the part worth reading is four replies deep in a thread I was never going to open.
It's not that I don't have the time to dig for the takeaways myself. It's that when I do have time like that, I don't want to spend it scrolling. So I've kept to just the headlines, which is leaving a lot of important context unread.
What Is This, Exactly?
Hacker News Radio reads the day's top stories, the posts themselves and the comments, and turns them into a podcast. Plain text in, a produced show out, voiced by Deepgram Flux TTS. It runs on a cron at 3am Pacific, writes the episode to a Fly volume, and publishes a real RSS feed with chapters and a transcript.
Each episode is five or six minutes, just enough time for me to get my coffee brewing before the day really starts, and it's yesterday's news by the time I hear it.
That bit, yesterday's news, sounded like a problem until I started listening day to day. At the pace these topics actually move, a day behind actually gives commenters time to respond and build up information, and remember the "old way" was me just not reading anything at all. A win's a win.
Listen To It Yourself
Listen live on the show page, or add it to your podcast player of choice: Apple Podcasts, Overcast, Pocket Casts, Castro, AntennaPod, or copy the raw RSS feed if your player isn't listed.

Build Rules, Self Imposed
I built this for Hacker News specifically, but the rule I gave myself from the start was that I should be able to point this at an RSS feed and walk away with a produced podcast.
That means the control I get over how a listener hears a story is close to none: plain spoken text, no markdown, no SSML, no bracketed cues. Every character in that string gets read out loud.
Enter: Flux TTS.
Most text-to-speech gets expressive by being told what to do. You wrap the text in markup, mark the emphasis, set the pauses, and hand the model a set of instructions alongside the words. Flux TTS does not need that layer. It takes the delivery from the writing itself, the punctuation and the sentence shape and what the sentence is actually saying, and that is the whole reason this project works. If I had to hand-mark every line, I could not point it at a feed and walk away.
Each segment is also one stateless call, so anything that sounds like a conversation is coming from the script and the gaps between lines, not from the model remembering what it just said.
That's the batch path, not a Flux TTS limit.
Cross-turn context is a real Flux TTS feature, it just lives on the streaming transport. Batch is stateless request-response by design. I took batch because the show is produced hours before anyone listens, so I have no latency budget to worry about.
Ok, small caveat to the "text in, every character out" mantra: I did end up creating a small dictionary file for text expansion. The single entry in there right now is to convert "HN" to "Hacker News".
Keeping Things Simple
Flux TTS gives you real control over how a line comes out. I turned almost none of it on, and that was the point, so here is what was sitting there.
Speed. A speech-rate multiplier, 0.85 to 1.15 in 0.05 steps. This is how fast the voice talks inside a line, which is a different thing from the silence between lines. I spent a lot of time on the second one and never touched the first.
Expressivity. A dial from -2, calm, to 2, animated. Still in beta. I expected to need this one, because "read this like a person would" is exactly the problem, and it turned out the writing was doing that work already.
Encoding, container, sample rate. The output format. I did set these, but for plumbing rather than taste: raw samples out, no header, so stitching an episode is concatenating bytes. Nothing about how it sounds.
Which voice. Thirty-six of them. This one I use constantly, and it is the only knob the show actually leans on.
There are more, and the voice controls docs are the place to look rather than this post.
So: two real taste knobs, and not one segment across nineteen published episodes has a value set on either. They were sitting right there and I still did not need them. The one place I did turn dials is music, and that happens after Flux TTS hands the audio back. Bed and cue levels are config, not TTS parameters.
Prototype Silence
Let's talk about the actual build process. After an initial pass working with Claude and other AI tooling, I got a basic version of the podcast generated, deployed it, and subscribed to the podcast in a podcast player (I used Overcast) and... nothing.
The generated RSS feed itself was valid and would load within podcast applications, but the audio itself just would not play.
So what was broken? It took a while to diagnose the real problem because nothing in the chain itself was broken. Turns out, Flux TTS does not hand back an MP3 at all.
It returns raw sixteen-bit PCM at 24 kHz, because (duh) that is what I asked for. Pulse-code modulation, meaning uncompressed audio samples and nothing else. Asking for it headerless is what makes stitching the episode a matter of concatenating bytes. The MP3 gets made afterward by ffmpeg.
That is where it goes wrong, and it took me longer than I want to admit to see it. MP3 is not one format. There are two flavors, and the sample rate picks which one you get, not you. 32, 44.1, and 48 kHz give you the common one. 24 kHz does not, so the encoder silently hands back the other.
Both are legal MP3s. Podcast players only reliably play the common one. So every episode I published was a real MP3 that players would list, show cover art for, and then refuse to open.
One ffmpeg flag fixes it: resample to 44.1 kHz on the way out. Same bitrate, same file size, different flavor.
Who's Up First
The second issue I did not find with a debugger. I found it by listening.
The first iteration episodes had a couple of commenters reacting to the story, and I wanted them to sound like different people. They were not.
Guest voices were assigned by the order they appeared: the first commenter was always the same voice, then the second, and so on. Flux TTS has a library of 36 voices to choose from, and right out of the gate I would never get the chance to listen to more than a small handful of them.
To fix this, I pivoted the show format to be more of a co-hosted show. Alexis is the showrunner, and each episode now runs with a selected co-host. This keeps the show cast down to just two voices per episode, which is easier to follow in a conversational format, and gives longer-term listeners exposure to the full library of voices Flux TTS has to offer.
Tips for Tuning
This bit is for the devs, or frankly, make sure your agent reads this bit, when building your own show.
If you build one of these, do these two things first.
-
Cache the rawest thing you have. This pipeline caches raw per-segment PCM, before pacing and before music, and that one decision is why I could render six full comparison versions, two episodes through three pacing policies, plus a four-level music sweep, without spending a single TTS call. Caching the finished show instead would have put a price on every taste question, and I would have asked fewer of them.
-
And this took me a bit longer than I had expected: measure before you tune. The show originally sounded like a monologue. The gap between segments was a configured 0.45 seconds, so that is the number I spent my time turning. Bluntly, it was the wrong number. Flux TTS bakes about a quarter second of its own silence onto the tail of every segment and almost none onto the front, so a configured 0.45 second gap was landing at roughly 0.73 seconds of real dead air. Real conversational turns sit near 0.20 seconds, which is nowhere near where I was.
Run It Yourself
Clone it and follow the README. You need uv, ffmpeg, and a Deepgram API key.
An Anthropic key is optional, and it changes who writes the show rather than whether it runs. make episode defaults to the deterministic writer: it assembles the script from the summaries pulled out of each linked article, every TTS call is real, and you get a complete episode that sounds a little like a template. --writer claude is the one that writes the show on the feed I linked.
Deepgram gives new accounts $200 in credit and does not ask for a card, which is up to about 950 episodes at the five or six minutes I have been making.
Clone it and let it build you a show. One episode is 16 to 29 TTS requests, so you are well past your first call before you have listened to the thing you made.
The whole thing is on GitHub.