POLARIS / A RESEARCH JOURNEY

Our path through artificial intelligence.

We began by building local models around market data. We then explored language-model APIs and chart images. Here we explain what we worked on, why the questions changed and what each approach added to our understanding.

We tell the story of our development work here. The chapters describe methods and experience rather than measured trading outcomes.

What did “building our own AI” actually mean?

We began with the idea of building a local intelligence layer. The first job was to turn that ambition into questions a model could actually learn from data.

At the beginning, we wanted to prepare market data, train models ourselves and obtain their outputs in our own Python environment. By “our own AI,” we meant specialised models for numerical market questions. We were not trying to train a general-purpose language model from scratch; that would involve a very different scale of data and computing.

The practical work soon made the question more specific. What should a model see: one candle, a sequence of observations or a table of derived features? What should it estimate: a direction, a numerical range or something else? We had to choose the targets, arrange the inputs and keep the transformations consistent before a saved model could be useful.

We explored several branches from that starting point. LSTMs let us work with sequences, tree-based models with engineered features, and Prophet with a time-series formulation. Later we used provider APIs to send context to already-trained language models, and explored chart images through local computer vision. These branches gave us different tools for different questions.

What stayed with us was the importance of describing the task clearly. “Local” tells you where a calculation runs. It does not tell you whether the model learns from numbers, interprets text or detects objects in an image. We keep those distinctions visible throughout the chapters that follow.

Workflow / 01

Three routes into AI

Separate approaches, not a single combined model.

Teaching a model to read a sequence, not one candle

We used LSTM models to work with a short history of observations and describe the output as downward, neutral or upward states.

We wanted the model to see how measurements developed, rather than treating the latest candle as an isolated moment. That led us to LSTMs, which process observations in order and carry an internal state through the sequence. We prepared rolling windows so each input included a short history.

In the archived classification models, a window contains 30 observations with 689 features in each observation. Those are dimensions of the historical research models, not published trading presets. The preparation involved keeping the columns in the same order and applying the fitted scaling that belonged to the model.

We also had to define what the three target classes meant. The intended reading was downward, neutral or upward, but the observation horizon and the rules used to create the labels gave those words their actual meaning. Changing a label definition would change the question even if the network architecture stayed the same.

We worked with a compact LSTM and a larger stacked version. Each returned three class scores. We treated those scores as the model's description of the input, not a complete future price path. This work made the companion feature list, scaler and target definition feel just as important as the saved network.

Workflow / 02

From history to a state

Class scores, not trade instructions.

Beyond direction: describing the possible extent of a move

We also asked for numerical outputs, so we could study the possible extent of a move instead of expressing every answer as a direction.

A directional label left part of our question unanswered. Two upward moves can have very different sizes and different movements along the way. We explored regression to express quantities, with two output values rather than three class scores.

The archived variants use 240 features and windows of 30 or 60 observations, with stacked LSTMs and attention-style components. We used these structures to work with different amounts of history and ways of weighting internal representations. We still had to explain the meaning and unit of each output; a more elaborate network does not do that for us.

We worked on transformations on both sides. The input had to match the training scale, and the output had to be returned to its intended units. Separate inference scripts also used a saved feature selector with a single-step LSTM input. We keep that formulation distinct from the 30/60-step models. In the working reports, observed maximum and minimum differences sit beside their predicted counterparts.

Range work also made us pay attention to order. A path can reach its low before its high, or its high before its low, while sharing the same two extreme values. We use the illustration alongside this chapter to make that distinction visible. Estimating two limits does not reconstruct everything that happens between them.

Workflow / 03

From history to two values

Two extrema do not reconstruct the price path.

Why the research did not stop at neural networks

We brought XGBoost, CatBoost and quantile outputs into the research so that different types of numerical questions had different tools.

Our work did not stay with neural networks. We also used tree-based model families, including XGBoost and CatBoost. Their way of partitioning feature values differs from an LSTM processing an ordered history, even when both begin with measurements from the same market.

We worked on serving scripts that loaded several model families and exposed their outputs through a common interface. That made it possible to inspect them within a shared workflow. We kept a separate question in view: making outputs available does not, on its own, establish how they should be combined or whether an ensemble adds value.

Quantiles gave us another way to describe an output. One archived interface used the 20th, 50th and 80th percentiles: a lower estimate, the median and an upper estimate. The interval from 20% to 80% has nominal central coverage of 60% if those quantiles are calibrated. We are careful not to describe that as an 80% confidence guarantee.

This broadened the language we could use in research. A class names a state, a point estimate gives one quantity and a quantile band describes positions in a conditional distribution. We could then choose an output according to the question, rather than treating every model score as interchangeable.

References behind the methodXGBoost · quantile regression
Workflow / 04

Different models, different outputs

Parallel routes do not imply automatic consensus.

Prophet and ARIMAX: another language for time

We explored time-series approaches alongside feature-heavy models, which made us think more explicitly about time and forecast horizons.

After working with feature matrices and rolling windows, we also explored a more direct time-series formulation. Our Prophet script used timestamped XAUUSD observations at 15-minute intervals and requested five future observations. The emphasis shifted toward the evolution of the series itself.

Prophet gave us a framework with a trend and optional seasonal components. It also made the horizon easy to explain: five 15-minute steps correspond to 75 minutes during a continuous trading session. We still need to account for closures and missing observations; five rows do not always mean uninterrupted clock time.

Our working archive also includes an ARIMAX-labelled report with timestamps and observed and predicted maximum/minimum differences. ARIMAX combines autoregressive structure with external explanatory variables. The report preserves that strand of work, although it does not reveal every modelling choice behind it. We do not reconstruct those missing choices from the filename.

Moving between these approaches changed the questions we asked. With a sequence network, we choose the ordered feature history to present. With a time-series model, temporal structure is built into the formulation. Both require us to be explicit about when an observation belongs and what “ahead” means.

Workflow / 05

Time gives the forecast its meaning

Five M15 steps; the ARIMAX report is separate.

From training a model to preparing a conversation with one

Using hosted language models shifted our work toward collecting context, framing a request and keeping the returned answer understandable.

When we began using language-model APIs, we were working with models that providers had already trained. Our contribution moved to the surrounding workflow: gathering information, deciding what to include and asking a question in a form we could record and review.

Our Python pipeline assembled MT5 candles from M5, M15 and M30, a current tick and timestamped public analyst posts. We treated prices as observations and commentary as interpretations. Keeping their origins and times mattered because several confident statements can still refer to different moments or disagree about the same market.

We asked for a JSON-formatted plan containing scenarios, conditions, invalidation and a no-trade option. That structure made the response easier to pass to software and inspect later. Requesting JSON, however, did not settle whether the answer's contents were valid. The research script saved and printed the response; it was a context-and-response workflow, not evidence of a complete execution controller.

This route taught us something different from local model training. We did not need to claim ownership of the underlying language model to do substantial work around it. Choosing the evidence, framing the request and preserving the answer were important parts of using the service meaningfully.

Workflow / 06

From observations to a response

Context and response—not an execution controller.

What if the model looked at the chart itself?

We explored a local YOLOv8 route to ask whether the chart image itself could become the input, rather than only numbers or text.

When we look at a chart, we notice spatial relationships: clusters of candles, developing moves and marked structures. We wanted to explore that representation directly, so our work included a local YOLOv8-based route for chart images.

That changed the task. An object detector locates and classifies visible objects, usually with boxes and scores. We had to think about which chart structures could be labelled consistently and what would count as an example. Detecting a visible shape is different from explaining a screenshot in language, and different again from forecasting a future price.

The image itself also raised questions. Zoom, cropping, candle width, colours and added annotations all change the pixels. We therefore had to distinguish display choices from the market information we wanted to represent. Examples where the sought structure is absent matter too; a detector should not be expected to find a shape in every chart.

Finally, a detector reports locations in pixels while trading work refers to price and time. Relating those two descriptions requires the chart scale and coordinate system. This is why we see vision as its own branch of the journey, with its own preparation and interpretation work.

References behind the methodUltralytics · object detection
Workflow / 07

From an image to visible objects

Illustrative boxes, not a real detector result.

The bridge between MT5 and Python

We worked on the bridge that let MT5 observations and Python models exchange information with a shared meaning.

Our market observations and trading-system logic lived on the MT5 side, while Python provided the model libraries and research tools. Connecting them meant more than passing a list of numbers. We needed to agree on what crossed the boundary and what the returned values meant.

We developed several inference-service variants that loaded combinations of LSTM, tree and quantile models. They prepared incoming data and returned model outputs through a common interface. This was a separate piece of engineering from training: each model expected a particular arrangement of observations and features.

The companion files became important at that point. A feature list preserved column meaning, an input scaler preserved the training transformation and an output scaler restored regression units. We had to interpret model version, sequence length, instrument and timeframe together. Receiving a response only proved that a request had been answered, not that those meanings matched.

This work made us more attentive to the complete model bundle. We explain its public architecture here while keeping service addresses, account connections and operating settings private. The useful part of the story is how the pieces fit, including the information needed to interpret a new observation consistently.

Workflow / 08

One request, one response

The service and its model bundle stay distinct.

The less visible work: datasets, selectors and working reports

We describe the dataset preparation, small utilities and working reports that supported the more visible model work.

A saved network or a prediction is easy to point to. Much of our effort sat before and after it. We prepared wide feature tables with time and OHLC fields, separate target tables, extraction utilities and prediction exports so the different research branches could be worked on coherently.

In the prediction scripts, one path reloaded a saved feature selector before shaping LSTM input; another reloaded the named feature list for XGBoost. Both kept time and a reference field alongside two predicted outputs. A separate extraction utility selected a bounded part of a dataset for a working file. These smaller tools helped us keep track of what each calculation was using.

We used CSV and Excel exports to inspect the work outside the training code. Some contained timestamped predictions; others placed observed and predicted maximum/minimum differences next to one another. Even file encoding needed attention because the working data included both UTF-16 and UTF-8. We discuss the structure here without publishing the private rows or inferring success rates from file labels.

Looking across the archive, we see a common thread: keeping enough context to understand where an answer came from. With a local model, that includes the data and transformations. With a hosted model, it includes the sources, prompt and returned response. The tools differ, but both ask us to preserve the reasoning around the output.

Workflow / 09

Keep the research trail

Keep the inputs beside the output.

We kept returning to the same question: what does this output actually mean?

Working with numbers, language and images gave us different ways to approach a market. It also made us pay more attention to the information behind an answer: where it came from, when it was available and how another part of the system would interpret it. That is the thread connecting the work in this notebook.

Talk about the work
Contact POLARIS on WhatsApp