Séminaires RALI-OLST-ILFC

Les mercredis à 11 h 30 (heure de Montréal), nous tenons un séminaire d'une heure portant sur un sujet du traitement des langues ou de linguistique. Il est typiquement offert en mode hybride (présentiel & vidéoconférence). Une fois par mois, le séminaire est organisé par le groupe de recherche français Linguistique Informatique, Formelle et de Terrain.

Sentiment-Augmented Deep Reinforcement Learning for Active Trading - An Alpha Reward Approach

Andrei Neagu (andrei (point) neagu <at> concordia (point) ca)

Concordia University

Le mercredi 4 novembre 2026 à 11 h 30

Salle 3195, Pavillon André-Aisenstadt — En présentiel, avec diffusion simultanée sur Zoom


We present our system for Task 3 of the CLEF 2026 FinMMEval Lab, which requires participants to submit a daily trading decision (long, flat, or short) for Bitcoin (BTC) and Tesla (TSLA) based on news articles and historical market pricing data. We frame the problem as a discrete-action Markov Decision Process and train four Deep Reinforcement Learning (DRL) algorithms: Policy Gradient (PG), Proximal Policy Optimization (PPO), Deep Q-learning (DQL), and Deep Deterministic Policy Gradient (DDPG), using a rich feature set that combines technical indicators (EMA, RSI, MACD, BollingerB, volume change), cyclical date encodings, and daily sentiment scores derived from news articles using LLaMA 3.2 1B. To reduce overfitting and align the training objective with the goal of outperforming a buy-and-hold baseline, we introduce an alpha reward that replaces the raw log-return with excess return over the market, and we randomize episode start days during training. Hyperparameters are optimized over 180 trials per algorithm-asset pair using Ray Tune, with model selection based on validation Sharpe ratio (SR) and early stopping. Evaluation on the CLEF Task 3 test set demonstrates that DDPG yields the best performance across both assets. Note, however, that DQL was selected a priori for the live competition endpoint based on its highest validation Sharpe ratio; the endpoint model was deliberately selected blind to the test period to avoid selection bias. For TSLA, DDPG and DQL achieved cumulative returns of 54.96% and 52.62% respectively, substantially beating the 16.45% buy-and-hold baseline. On the more challenging BTC test set, DDPG mitigated severe market losses, achieving a positive return of 1.58% against a baseline decline of -34.27%. Results reveal a pronounced validation-to-test generalization gap, pointing to the difficulty of adapting policies trained in a bull market validation period to bear market test conditions.



Joignez-vous à nous avec Zoom à l'aide de cette URL.


Suivez ce lien pour vous inscrire à la liste de diffusion RALI-OLST.
http://rali.iro.umontreal.ca/rali/?q=fr/node/1631

Liste de tous les séminaires pour l'année :

1991 1992 1993 1994 1995 1997 1998 1999 2000 2001 2002 2003 2004 2005 2006 2007 2008 2009 2010 2011 2012 2013 2014 2015 2016 2017 2018 2019 2020 2021 2022 2023 2024 2025 2026