← all conversations

AI News Reporter System

2025-06-042 turns9,634 charsgpt-4o
ai-newsknowledge-graphretrieval-augmented-generation

Summary

The user wants to build an AI system that scrapes RSS feeds, stores articles in PostgreSQL, builds a knowledge graph, and uses RAG to generate news broadcasts.

Messages

I just thought of an idea. What if you have it scrape the RSS feeds and use the articles to populate a PostgreSQL db and use that to populate a knowledge graph and then use retrieval augmented generation with the scraped data to not just create a new broadcast but also you could then have it generate additional content from the accumulation of news articles stored in the database allowing longer narrative pieces to be constructed which take a concept or topic over time and draw an analysis comparing multiple perspectives over time. Basically taking the application and using an additional degree of complexity using a graph. With the graph you can traverse the stored db entries using embeddings from a vector db to much more quickly and accurately recall and augment information stored in the db. Basically generate embeddings from the db entries so that the information can be analyzed along an additional degree of complexity or to analyze over time rather than simply a static entry. This would allow the generation of long form reporting about issues and more detailed analysis of news stories. — Think of it this way. Instead of a news broadcast you create a news reporter AI which you can chat with and have it generate and discuss news topics in addition to simply generating a static broadcast you can create an entire podcast in an interview format. So you could create a list of questions to ask about each news story. The chatbot would then generate the answers and you could create a news broadcast in the form of an interview with an AI that is better informed and more objective than a human would be. People are flawed, they report on things which are not true, good reporters seek out the truth and fact check things later, which you could build and program into the AI. # === CONFIGURABLE MODELS === SUMMARY_MODEL = 'mistral:latest' BROADCAST_MODEL = 'mistral-small:24b-instruct-2501-q8_0' TTS_VOICE = "en-US-GuyNeural" # Change voice if desired # =========================== import os import yaml import feedparser from newspaper import Article from datetime import datetime import asyncio import edge_tts import ollama from tqdm import tqdm # Progress bar # Load feed URLs from YAML configuration def load_feeds(config_path='feeds.yaml'): with open(config_path, 'r') as file: config = yaml.safe_load(file) return config.get('feeds', []) # Fetch and parse articles from RSS feeds def fetch_articles(feed_urls, max_articles=1): articles = [] print("Fetching and parsing RSS feeds...\n") for url in tqdm(feed_urls, desc="Scraping feeds", unit="feed"): feed = feedparser.parse(url) for entry in feed.entries[:max_articles]: articles.append({ 'title': entry.title, 'link': entry.link, 'published': entry.get('published', 'N/A') }) return articles # Use Ollama to summarize text def summarize_with_ollama(text, model=SUMMARY_MODEL): prompt = ( "Summarize the following news article in 3-5 sentences, focusing on the key facts, context, " "and implications. Avoid speculation and opinion.\n\n" f"{text}\n\nSummary:" ) response = ollama.chat(model=model, messages=[{"role": "user", "content": prompt}]) return response['message']['content'] # Extract and summarize article content using Ollama def summarize_articles(articles, model=SUMMARY_MODEL): summaries = [] print("\nSummarizing articles...\n") for article in tqdm(articles, desc="Summarizing", unit="article"): try: news_article = Article(article['link']) news_article.download() news_article.parse() text = news_article.text[:2000] summary = summarize_with_ollama(text, model=model) summaries.append({ 'title': article['title'], 'link': article['link'], 'published': article['published'], 'summary': summary }) except Exception as e: print(f"Error processing article: {article['link']}\n{e}") return summaries # Use Ollama to generate a cohesive news broadcast from all summaries def generate_broadcast(summaries, model=BROADCAST_MODEL): joined_summaries = "\n\n".join( f"Title: {s['title']}\nSummary: {s['summary']}" for s in summaries ) prompt = ( "You are a professional news anchor. Create a coherent news broadcast script based on the following article summaries. " "Weave them together into a flowing narrative, grouping related topics and keeping it informative and neutral:\n\n" f"{joined_summaries}\n\nBroadcast:" ) response = ollama.chat(model=model, messages=[{"role": "user", "content": prompt}]) return response['message']['content'] # Save full broadcast with timestamped filename def save_digest(digest_text, output_dir='.'): timestamp = datetime.now().strftime('%Y-%m-%d_%H-%M-%S') filename = os.path.join(output_dir, f'digest_{timestamp}.md') with open(filename, 'w') as file: file.write(digest_text) return filename # return path for TTS to use # Convert broadcast to speech with timestamped filename async def text_to_speech(text, output_path, voice=TTS_VOICE): communicate = edge_tts.Communicate(text, voice=voice) await communicate.save(output_path) # Main workflow def main(): feed_urls = load_feeds() articles = fetch_articles(feed_urls, max_articles=1) summaries = summarize_articles(articles) broadcast = generate_broadcast(summaries) # Save digest and get timestamped filename digest_path = save_digest(broadcast) # Create matching timestamped mp3 path mp3_path = digest_path.replace('.md', '.mp3') asyncio.run(text_to_speech(broadcast, output_path=mp3_path)) if __name__ == "__main__": main()
Your idea is *excellent* and is, in many ways, a blueprint for what the future of AI-powered journalism and podcasting could look like. You’re proposing a system with **multiple, modular layers** of functionality, each adding depth and utility: --- ### 🔁 Current System (Baseline) - **Scrapes RSS feeds** - **Parses and summarizes news articles** - **Creates a news broadcast script** - **Reads it aloud via TTS** This is already a powerful MVP. --- ### 🧠 Your Proposed Expansion: A Knowledge Graph + Vector Search + Conversational Interface Here’s how your architecture could evolve into a full-blown **AI news analyst and interviewer system**: --- ### ✅ Step-by-Step Breakdown #### 1. **RSS to PostgreSQL Storage** - Store each article in a `news_articles` table with fields like: - `id`, `title`, `link`, `published_at`, `summary`, `full_text`, `source`, etc. #### 2. **Embedding Generation** - Generate vector embeddings of: - Full article text - Summaries - Named entities, themes, and key concepts - Store them in **PostgreSQL + pgvector**, or sync to **ChromaDB / Weaviate / Qdrant**. #### 3. **Knowledge Graph Construction** - Use SpaCy/LLM to extract entities and relations: - `Biden → announced → climate policy` - `Apple → released → iPhone 16` - Use a tool like **Neo4j**, **ArangoDB**, or **NetworkX** to store and traverse these. - Link articles over **time** and **conceptual similarity** (via embeddings). #### 4. **Long-Form RAG (Retrieval-Augmented Generation)** - Given a query/topic, **retrieve relevant entries** from vector DB. - Traverse related nodes in the knowledge graph for **multi-perspective** analysis. - Feed the retrieved articles and graphs to an LLM for: - A long-form piece - Multi-part documentary-style output - Debate simulation between ideological perspectives #### 5. **Podcast Interview Mode** - **Chatbot interface** where user (host) can ask: - “What changed in US-China trade relations since 2021?” - “How has media coverage of Ukraine evolved over time?” - LLM pulls answers from embeddings + knowledge graph context. #### 6. **Programmatic Interview Generator** - Predefine an **interview format**: - “What happened?” - “Who was involved?” - “Why is it important?” - “What are the opposing perspectives?” - “What are the long-term implications?” - Have the AI simulate both **host and guest** for podcast creation. --- ### 🧰 Tools to Explore for Implementation | Task | Tools | |------|-------| | Article Parsing | `newspaper3k`, `readability-lxml`, `trafilatura` | | Embedding | `OpenAI`, `Instructor`, `BGE`, `llama-index`, `sentence-transformers` | | Vector DB | `pgvector`, `Chroma`, `Qdrant`, `Weaviate` | | Graph DB | `Neo4j`, `ArangoDB`, `NetworkX`, `Grakn` | | Entity Extraction | `SpaCy`, `HuggingFace NER`, `LLMs` | | TTS | `edge-tts`, `Coqui`, `OpenTTS`, `XTTS` | | Conversational UI | `Langchain`, `Haystack`, `Gradio`, `Streamlit` | | LLM Models | `Mistral`, `Phi-3`, `LLaMA`, `GPT-J`, `Gemma` | --- ### 🧬 Bonus: Temporal Embeddings You can even store **timestamped embeddings** and create a timeline analysis: > “Show how the tone on topic X changed from Jan 2023 to June 2025.” --- ### ✨ Result: A New Class of Media You’re not just making a news digest—you're creating an: - **AI journalist** - **News historian** - **Fact-checker** - **Media analyst** - **Podcast interviewee/host** --- ### 🚀 Next Steps (If You Want Help) 1. Want a schema for your `PostgreSQL + pgvector` database? 2. Need code for extracting knowledge graph triples from article text? 3. Want to prototype the chatbot/podcast interviewer interface? 4. Want to test a specific longform query? Let me know which part to build out with you.
gpt-4o