ETFs portfolio
Summary
A personal portfolio manager that pulls live holdings from brokers, optimizes allocation with mean-variance analysis, and charts growth and drawdown against benchmarks.
The Problem
Investing can take many forms, and the stock market is one of them. I chose to include it as one part of diversifying my investments. To invest thoughtfully, I first needed to understand how the market works and the mathematical theory behind it. The next challenge was putting that understanding into practice by building a portfolio of my own. Given the current market, an ETF-based portfolio made sense to me, so I built one.
The Solution
I started by reading the basic mathematical theory and building an intuitive understanding of how it works. But I also needed to understand the practical decisions involved in building a portfolio:
- What exactly is an ETF, what are its main costs, and how do I choose one?
- Why use different ETFs, and how do I choose them to diversify my portfolio?
- How does tax work for ETFs in this context?
- How do I choose portfolio weights, and which optimizers should I use?
- How much historical data do those optimizers need?
- Once I have a portfolio, how should I monitor it? Which indices should I compare it with?
- How do I check that the whole data pipeline works correctly? Can I cross-check it against other sources at different stages?
Each of these questions led to more mathematical and programming questions. I read different resources to understand the subject and what I was getting into. The project also led me to build a Python tool to manage and analyze my own ETF portfolio.
I started by gathering ETF data. I used SQLite as the database and SQLAlchemy to work with it from Python. This approach can be slower in some places than writing SQL directly, but I found the code easier to understand and maintain. For this project, that trade-off was fine: parsing data, not storing it, was the main bottleneck.
At first, I wrote a parser for Interactive Brokers data using the ib_async library, which I found straightforward to work with. After saving the data, I noticed missing records on days when I expected data to be available. I traced the gaps to data from Interactive Brokers through SMART, which did not always include the records I needed. I added Yahoo Finance, accessed through the yfinance library, as another source to fill those gaps.
I then ran additional checks across the parsed data to look for anomalies. Most of the data checked out, and I could account for the irregularities I found. With the data collected and checked, I moved on to analysis.
I tried several portfolio optimization approaches: mean-variance, Conditional Value at Risk, Risk-Averse Hierarchical Risk Parity, Hierarchical Risk Parity, and Minimax Loss. Mean-variance optimization focuses on expected return. I used PyPortfolioOpt and Riskfolio-Lib to build and compare the models, and used block bootstrapping to generate additional samples from historical data. I used as much historical data as I could so the models had more data to work with.
When choosing ETFs, I aimed to include exposure to European, US, and Asian markets, along with funds with lower and higher volatility. Most of my choices were global ETFs, which already provided broad geographic diversification. I generated different portfolio options along the efficient frontier, then combined the variants I preferred to diversify across the ETFs. The final portfolio gave equal weight to the mean-variance and Minimax Loss approaches.
Once I had a portfolio, I tested it against a benchmark across different historical start dates to see how it would have performed if I had invested at those times. I compared its performance measures and drawdowns. In the periods I tested, it outperformed the benchmark without experiencing very large drawdowns.
I also verified the process at each stage using several tools that I can't name for personal reasons. I checked the work as I went and built the portfolio step by step.
The chart compares my portfolio with the ticker symbols of the ETFs it contains and with the benchmark index, shown in pink. Some ETFs do not have historical data for the full period, so I assume a 0% return for them until their data becomes available. After that, their actual returns are included using their assigned portfolio weights. I ran the optimizers only over periods for which historical data was available for every ETF.

The drawdowns differ quite a lot. The earlier part of the chart is less representative because several ETFs lack historical data, so their returns are assumed to be 0%, which softens the apparent drop. The comparison is clearer over the most recent ten years: the pink line falls further than the black portfolio line. After repeating the comparisons and looking at the performance figures, I chose this portfolio because it seemed like a good fit for me. I’ll only know how well it works for me over the long term after living with it for several years.
