diff --git a/agent.md b/agent.md new file mode 100644 index 0000000..a695a94 --- /dev/null +++ b/agent.md @@ -0,0 +1,298 @@ +# StockDocs Project Overview + +This document provides a comprehensive overview of the StockDocs repository, which contains multiple interconnected projects for processing, scraping, and serving financial articles and embeddings. + +## Repository Structure + +The repository contains 5 main projects: + +1. **ai_processor/** - AI processing component +2. **articleServer/** - Article serving component +3. **embedding/** - Embedding functionality +4. **MCPServer/** - MCP server component +5. **scraper/** - Web scraping component + +## Project Details + +### 1. ai_processor/ +The AI processing component handles artificial intelligence operations for processing articles and generating insights. + +**Key Files:** +- `ai_processor.py` - Main AI processing logic +- `requirements.txt` - Python dependencies +- `Dockerfile` - Container configuration + +**Purpose:** Processes articles using AI models to extract key information, generate summaries, and create embeddings. + +**Detailed Structure:** +``` +ai_processor/ +├── app.py # Main AI processing application +├── config.py # Configuration settings +├── requirements.txt # Python dependencies +├── Dockerfile # Docker configuration +├── README.md # This file +├── models/ # Machine learning models and NLP components +│ ├── __init__.py +│ ├── sentiment_analyzer.py # Sentiment analysis module +│ ├── topic_classifier.py # News categorization module +│ └── entity_extractor.py # Named entity recognition +├── processors/ # Article processing pipelines +│ ├── __init__.py +│ ├── text_processor.py # Text cleaning and preprocessing +│ └── analysis_pipeline.py # Full analysis pipeline +└── data/ # Processed data storage + ├── insights/ + └── reports/ +``` + +**Endpoints:** +- POST `/api/analyze/article` - Analyze a single article for insights +- POST `/api/analyze/batch` - Process multiple articles in batch mode +- GET `/api/analyze/status/{task_id}` - Check processing status +- GET `/api/insights/latest` - Get latest analysis insights +- GET `/api/insights/articles/{article_path}` - Get insights for specific article +- GET `/api/reports/generate` - Generate comprehensive market analysis report +- GET `/api/models` - List available AI models +- POST `/api/models/update` - Update or retrain models with new data + +### 2. articleServer/ +The article serving component provides an API for accessing processed articles. + +**Key Files:** +- `app.py` - Main Flask application +- `run_server.py` - Server startup script +- `requirements.txt` - Python dependencies +- `Dockerfile` - Container configuration + +**Purpose:** Exposes processed articles through a REST API for client applications to consume. + +**Detailed Structure:** +``` +articleServer/ +├── app.py # Main Flask application +├── run_server.py # Server startup script +├── requirements.txt # Python dependencies +├── Dockerfile # Docker configuration +├── README.md # This file +└── templates/ # HTML templates (if any) +``` + +**Endpoints:** +- GET `/articles` - Get articles within time range +- GET `/article/content` - Get full article content by path +- GET `/outlets` - List all available news outlets +- GET `/health` - Health check endpoint + +### 3. embedding/ +The embedding functionality handles vector embeddings for articles and documents. + +**Key Files:** +- `embedder.py` - Embedding generation logic +- `requirements.txt` - Python dependencies +- `Dockerfile` - Container configuration + +**Purpose:** Converts articles into vector embeddings for semantic search and similarity operations. + +**Detailed Structure:** +``` +embedding/ +├── app.py # Main embedding application +├── config.py # Configuration settings +├── requirements.txt # Python dependencies +├── Dockerfile # Docker configuration +├── README.md # This file +├── models/ # Pre-trained embedding models +│ ├── __init__.py +│ ├── sentence_transformer.py # Sentence transformer implementation +│ └── model_loader.py # Model loading utilities +├── processors/ # Text processing pipeline +│ ├── __init__.py +│ ├── text_cleaner.py # Text cleaning and preprocessing +│ └── embedding_generator.py # Embedding generation +└── data/ # Processed embeddings storage + ├── cache/ + └── outputs/ +``` + +**Endpoints:** +- POST `/api/embeddings/generate` - Generate embeddings for text content +- POST `/api/embeddings/batch` - Generate embeddings for multiple texts in batch +- GET `/api/embeddings/similarity` - Calculate similarity between two pieces of text/content +- GET `/api/embeddings/status` - Check service status +- GET `/api/embeddings/models` - List available embedding models +- DELETE `/api/embeddings/cache/clear` - Clear the embedding cache + +### 4. MCPServer/ +The MCP (Model Context Protocol) server component provides external API access. + +**Key Files:** +- `server.py` - Main MCP server implementation +- `openapi.json` - API specification +- `requirements.txt` - Python dependencies +- `Dockerfile` - Container configuration + +**Purpose:** Exposes functionality through the Model Context Protocol for integration with other systems. + +**Detailed Structure:** +``` +MCPServer/ +├── server.py # Main Flask application +├── config.py # Configuration settings +├── requirements.txt # Python dependencies +├── Dockerfile # Docker configuration +├── README.md # This file +└── api/ # API endpoints and handlers + ├── __init__.py + ├── stock_analysis.py # Stock analysis functions + └── news_processing.py # News processing functions +``` + +**Endpoints:** +- GET `/api/stock/metrics` - Get financial metrics for a stock +- GET `/api/stock/history` - Get historical price data +- POST `/api/stock/analyze` - Perform comprehensive stock analysis +- GET `/api/news` - Retrieve news articles related to stocks +- GET `/api/news/outlets` - List available news sources +- POST `/api/news/process` - Process and categorize news content +- GET `/api/data/refresh` - Refresh data from sources +- GET `/api/status` - Server health check + +### 5. scraper/ +The web scraping component collects articles from various sources. + +**Key Files:** +- `scraper.py` - Main scraping logic +- `rss_feeds.json` - RSS feed configuration +- `requirements.txt` - Python dependencies +- `Dockerfile` - Container configuration + +**Purpose:** Collects financial articles from various sources including RSS feeds and web scraping. + +**Detailed Structure:** +``` +scraper/ +├── rss_feeds.json # Configuration file with RSS feed URLs +├── scraper.py # Main scraping logic +├── requirements.txt # Python dependencies +├── dockerfile # Docker configuration +├── dockerfile-selenium # Dockerfile for selenium-based scraping +├── articles/ # Directory where scraped articles are stored +│ ├── Reuters – Business News/ +│ │ ├── article1.txt +│ │ └── ... +│ ├── Associated Press – Business/ +│ │ ├── article1.txt +│ │ └── ... +│ └── ... +├── processed_articles_cache.json # Cache of already processed articles +└── pyvenv.cfg # Python virtual environment configuration +``` + +**Features:** +- RSS Feed Integration: Supports multiple financial news sources through RSS feeds +- Automated Scraping: Regularly fetches and processes new articles from configured feeds +- Structured Storage: Organizes articles in a directory structure by news outlet +- Duplicate Detection: Prevents re-processing of already collected articles +- Caching Mechanism: Maintains a cache of processed articles to optimize performance + +## Getting Started + +### Prerequisites +- Docker installed +- Python 3.8+ +- Git + +### Setup Instructions + +1. **Clone the repository:** + ```bash + git clone + cd StockDocs + ``` + +2. **Build and run containers:** + ```bash + docker-compose up --build + ``` + +3. **Project-specific setup:** + - Each project has its own `README.md` with detailed setup instructions + - Check individual project directories for specific requirements + +## Project Dependencies + +### Common Dependencies +- Python 3.8+ +- Docker +- Various Python packages (listed in requirements.txt files) + +### Inter-project Relationships +- `scraper/` feeds articles to `ai_processor/` +- `ai_processor/` generates embeddings that `embedding/` processes +- `articleServer/` serves articles processed by `ai_processor/` +- `MCPServer/` provides API access to all components + +## Development Workflow + +1. **Start all services:** + ```bash + docker-compose up --build + ``` + +2. **Work with individual projects:** + - Navigate to project directory + - Check `README.md` for specific instructions + - Make changes and rebuild as needed + +3. **Testing:** + - Each project includes its own testing setup + - Integration tests may be needed for cross-project functionality + +## API Endpoints + +### articleServer/ +- `/articles` - Get all articles +- `/articles/` - Get specific article +- `/search` - Search articles by query + +### MCPServer/ +- Exposes various endpoints through Model Context Protocol +- See `openapi.json` for complete specification + +## Configuration + +### Environment Variables +Each project may require specific environment variables. Check individual `README.md` files for details. + +### Data Storage +- Articles are stored in the scraper component +- Processed data flows through the ai_processor +- Embeddings are generated and stored in the embedding component + +## Troubleshooting + +### Common Issues +1. **Docker build failures:** Ensure Docker is running and check `Dockerfile` syntax +2. **Python dependency issues:** Run `pip install -r requirements.txt` in each project +3. **Port conflicts:** Check `docker-compose.yml` for port mappings +4. **Service startup issues:** Check individual project logs + +### Logs +- View logs with `docker-compose logs ` +- Check individual project logs for detailed error information + +## Contributing + +1. Fork the repository +2. Create feature branch +3. Make changes +4. Test thoroughly +5. Submit pull request + +## Support + +For issues or questions, please check: +- Individual project README.md files +- Docker logs for runtime errors +- GitHub issues for known problems \ No newline at end of file