6 ms·
if the data is not synthetic, how do you ensure that the LLM hasn't learnt about this data for example from training on the Financial Times.
by ak_111 1mo ago
if the data is not synthetic, how do you ensure that the LLM hasn't learnt about this data for example from training on the Financial Times.
- Mzzzzz 1mo agoWe do a 2 step anonymisation: 1. Mask all symbols, timestamps etc. So the agents cannot infer the assets/time periods. 2. Mathematically transform numerical values and returns. E.g. the market return targets are not the raw market returns, but neutralised and manipulated. So even the agents have certain bullish/bearish biases, it cannot make use of it, as we use the transformed values. In addition, we did not observe such behaviour in our traces. An example: https://hub.harborframework.com/jobs/af0299f9-a3bb-44ea-8ced-7b56682bf739 https://hub.harborframework.com/jobs/af0299f9-a3bb-44ea-8ced...
- ak_111 1mo agoah i thought so, interesting. I think the challenge is to do 2 while still keeping it realistic, which actually gets very close to synthetic data generation.
- RuiWang0811 1mo agowe do affine transformations of the data, so all return/ pnl measures are still the same as with untransformed data. The transformation doesn’t change the conditional distribution of the data, which is what alphas ultimately measure