4
5 years ago I trained a model on my old blog posts, the output was weirdly familiar
Back in 2019 I scraped about 400 posts from my LiveJournal and ran them through a basic LSTM just for fun. The thing that stuck with me wasn't the gibberish, it was how it picked up my tics like ending sentences with 'or whatever' and overusing the word 'literally'. Now everyone talks about fine-tuning LLMs on personal data, but most people skip the cleaning step. My raw posts had typos, meme references, and inside jokes that made the output useless for anything real. What I learned is that garbage in means you get a chatbot that sounds like you after three beers, not like a thoughtful version of you. Has anyone else tried feeding their old social media or journal entries into a modern model and actually gotten something useful out of it?
0 comments
Log in to join the discussion
Log In0 Comments
No comments yet
Be the first to share your thoughts on this discussion.