We are seeking an experienced Data Engineer to develop a robust real-time data collection and analysis system for options trading strategies. The ideal candidate will have a strong background in Apache Kafka and distributed systems, with specific experience in handling high-frequency financial data. This project involves integrating multiple data streams, processing real-time bid/ask quotes and trade executions, and building scalable pipelines to analyze the impact of market conditions on trading activities.
Responsibilities:
- Design, implement, and manage a scalable Apache Kafka cluster for real-time data ingestion and processing.
- Develop Kafka producers to capture live bid/ask prices and trade execution data from multiple sources.
- Create Kafka consumers and stream processing applications to analyze how bid/ask changes influence transactions.
- Integrate Kafka streams with external storage solutions (e.g., SQL/NoSQL databases, data lakes) for long-term storage and backtesting.
- Optimize data pipelines for high throughput and low latency to support high-frequency trading environments.
- Ensure synchronization and alignment of bid/ask data and trade executions for accurate analysis.
- Collaborate with stakeholders to understand trading strategies and data requirements.
- Implement monitoring and logging solutions to track system performance and ensure data integrity.
- Provide documentation and training for ongoing maintenance and system enhancements.
Requirements:
- Proven experience in setting up and managing Apache Kafka clusters in production environments.
- Proficiency in programming languages such as Java (essential for Kafka), Python, or Scala.
- Strong knowledge of real-time data processing and stream processing frameworks (e.g., Kafka Streams, Apache Flink, Apache Spark).
- Familiarity with SQL and NoSQL databases for storing and querying processed data.
- Solid understanding of distributed systems architecture, including data partitioning, replication, and fault tolerance.
- Experience in financial markets, particularly options trading and market data structures, is highly desirable.
- Ability to optimize systems for low latency and high throughput in high-frequency trading scenarios.
- Excellent problem-solving skills and the ability to troubleshoot complex issues in distributed systems.
- Strong communication skills, with the ability to explain technical concepts to non-technical stakeholders.
- Degree in Computer Science, Engineering, or a related field (preferred).
Preferred Skills:
- Experience with cloud platforms (AWS, Google Cloud, Azure) for deploying and scaling Kafka clusters and related infrastructure.
- Knowledge of Docker and Kubernetes for containerization and orchestration of services.
- Understanding of quantitative analysis and backtesting methodologies for financial data.
- Prior work on projects involving real-time data analysis and high-frequency trading.
Project Scope:
- Initial setup and configuration of Kafka cluster and data pipelines.
- Development and integration of data ingestion and processing applications.
- Ongoing optimization and maintenance of the system for performance and scalability.
- Potential for long-term collaboration on additional features and enhancements.
How to Apply:
Please provide:
- A brief overview of your experience with Kafka and real-time data systems.
- Examples of previous projects or case studies relevant to financial data processing.
- Your approach to handling real-time bid/ask and trade data for analysis.
- Your availability and estimated timeline for project milestones.
