Bir müşteri gelir, alacağı ürünleri listeler, sistem ona tek bir truck ve bir zaman dilimi ayırır. Sonraki müşteri geldiğinde rotası, o sırada depoda olan diğer truck'larla çakışmayacak şekilde planlanır. Depo hizmeti böyle, müşteri kuyruğu üzerinden kümülatif olarak optimize edilir.
Depoda aynı anda birden çok truck bulunur — ama bu çokluk tek bir müşterinin talebinden değil, arka arkaya gelen müşterilerden doğar. Müşteri sayfası bu yüzden ajan sayısı sormaz.
Aynı soruyu — "hangi rafa hangi sırayla gidilecek" — iki farklı yaklaşım çözüyor.
| ACO | Karınca Kolonisi Optimizasyonu. Eğitim gerektirmez, her koşuda sıfırdan arar. Sanal karıncalar farklı rotalar dener, iyi olanlar feromon bırakır, sonraki turlar o izleri takip eder. Her zaman kullanılabilir. |
|---|---|
| PPO | Pekiştirmeli öğrenme, politika tabanlı. Bir sinir ağı "bu durumda şu rafa git" demeyi önceden öğrenir. Kararı anında verir ama önce eğitilmesi gerekir. |
| DDDQN | Pekiştirmeli öğrenme, değer tabanlı. Her olası hareketin değerini tahmin eder, en yükseğini seçer. Eğitimi PPO'nun yaklaşık 9 katı sürer. |
RL seçenekleri gri görünüyorsa o depo için eğitilmiş model yok demektir. Model haritaya özeldir: aksiyon uzayı (seçilebilecek hareket sayısı) her haritada farklıdır — TOROS'ta 52, ANADOLU'da 368 — ve ağın çıkış katmanı tam bu genişlikte kurulur. Bir haritanın modeli başka haritada yüklenmez. ACO bu kısıttan muaftır.
Yüklemeyi truck değil, tavandaki vinç yapar. Truck yükleme şeridine park eder, vinç yatay eksende (şeride dik) ilerleyip rafa ulaşır ve yükü truck'ın üstüne bırakır. Bu yüzden truck'ın ürüne bitişik olması gerekmez — aynı satırda olması yeterlidir.
Karar: vincin taraması komşu yükleme şeridine
çarpınca durur; oradan sonrası bir sonraki gözün ve onun vincinin alanıdır.
Sınırsız erişim fiziksel olarak savunulabilirdi ama her ürün için onlarca
duruş noktası üretip arama maliyetini katlıyordu.
WH_CRANE_REACH_CELLS ile sınır konabilir (0 = şeride kadar).
Randevular üst üste binebilir: bir müşterinin işi sürerken bir başkası da başlayabilir. Kaç tanesinin aynı anda çalışabileceği haritanın kendi özelliğidir, global bir sabit değil.
Karar: kapasite = geçilebilir hücre sayısı / 10, 2 ile 30 arasında sınırlı, ve hiçbir zaman park yeri sayısını aşmaz. Neden park yeri değil de yoğunluk: ANADOLU'da 144 yükleme hücresi var ama 266 geçilebilir hücreye 144 truck koymak akış değil kilitlenmedir. Çok ajanlı yol planlamada ~%10 ajan yoğunluğunun üstünde tıkanıklık hızla artar.
Aynı sayı iki yeri birden belirler: kaç randevunun üst üste binebileceğini ve RL durum vektöründeki rota bandı sayısını. Böylece zamanlayıcının izin verdiği her truck, politikaya da görünür olur. Kapasite dolduğunda yeni randevu, yer açılan ilk ana kayar. Güncel değerler aşağıdaki tabloda.
Randevu, rota hesaplanmadan önce ayrılır — çünkü bir rota ancak içinde çalışacağı zaman diliminde anlamlıdır. Sıra şöyle:
t_end) güncellenir ve çizelge yeniden akıtılır.Rezervasyon zamanları gün başlangıcından saniye olarak saklanır ve planlamada yalnızca zaman penceresi çakışanlar yüklenir: sabah 8'de işi biten bir truck, öğleden sonraki bir müşteriyi engellemez. Kapsam aynı gün ve aynı depodur.
ACO'nun maliyeti sipariş büyüklüğünde karesel artar: her adımda kalan tüm ürünler, her ürün için de tüm olası duruş noktaları denenir. 1 ürünlü sipariş ~200 rota araması, 16 ürünlü sipariş ~27.000.
Sınırlı arama (varsayılan) toplam arama sayısını sabit tutar ve gerekirse tur/karınca sayısını düşürür. Tam arama istenen 10×10 aramayı aynen yapar; büyük siparişlerde süre sınırını aşabilir ve arama yarıda kesilir. Seçim kutusunun altındaki tahmin, ikisinin de kaç arama ve kaç dakika süreceğini sipariş değiştikçe canlı gösterir.
#12 gibi numaralarla ve renkle ayrılır.Video .webm (VP8) olarak üretilir. VP9
denendi ve aynı görüntü için 4× daha yavaştı; bu kareler düz
renkli olduğu için VP9'un sıkıştırma avantajı burada işe yaramıyor.
| Aksiyon uzayı | Modelin seçebileceği farklı hareket sayısı. Haritaya özeldir ve modelin hangi haritada çalışabileceğini belirler. |
|---|---|
| Episode | Bir işin baştan sona yapılması: girişten girip tüm siparişi toplayıp çıkışa varmak. |
| Zaman-indeksli A* | Yol bulma algoritması. Sıradan A* "hangi hücrelerden geçeyim" der; bu "ve ne zaman" boyutunu da ekler, böylece bekleme ve çakışma hesaba katılabilir. |
| Rezervasyon | "Şu hücre, şu zaman aralığında şu truck'a ayrılmış" kaydı. Çakışma önleme buradan çözülür. |
| Feromon | ACO'da karıncaların iyi rotalarda bıraktığı iz; sonraki turlarda o rotaların seçilme olasılığını artırır. |
Aşağıdaki tablodaki değerler bu sayfaya gömülü değil, çalışan sunucudan okunur; yani burada yazan kural ile sistemin uyguladığı kural ayrışamaz.
A customer arrives, lists the products they want, and the system allocates them one truck and one time slot. When the next customer arrives, their route is planned so that it does not collide with the other trucks in the warehouse at that time. The warehouse service is optimised this way — cumulatively, across the queue of customers.
There are several trucks in the warehouse at once, but that plurality comes from customers arriving one after another, not from a single customer's request. That is why the customer page never asks how many agents you want.
Two approaches answer the same question — "which rack, in which order".
| ACO | Ant Colony Optimisation. Needs no training; it searches from scratch on every run. Virtual ants try different routes, the good ones leave pheromone, and later rounds follow those trails. Always available. |
|---|---|
| PPO | Reinforcement learning, policy based. A neural network learns in advance to say "in this state, go to that rack". It decides instantly, but it has to be trained first. |
| DDDQN | Reinforcement learning, value based. Estimates the value of every possible move and takes the highest. Training costs roughly 9× what PPO does. |
If the RL options are greyed out, no model has been trained for that warehouse. A model is map specific: the action space (how many moves can be chosen) differs per map — 52 on TOROS, 368 on ANADOLU — and the network's output layer is built exactly that wide. One map's model will not load on another. ACO is free of this constraint.
Loading is done by the ceiling crane, not by the truck. The truck parks in a loading lane, the crane travels along the horizontal axis (perpendicular to the lane), reaches the rack and sets the load on top of the truck. So the truck does not have to be adjacent to the product — being in the same row is enough.
Decision: the crane's sweep stops when
it meets the neighbouring loading lane; beyond that belongs to the next
bay and its own crane. Unlimited reach would have been physically
defensible, but it produced dozens of stopping points per product and
multiplied the search cost. WH_CRANE_REACH_CELLS sets the
limit (0 = as far as the lane).
Bookings may overlap: one customer's job can still be running when another starts. How many can run at once is a property of the map, not a global constant.
Decision: capacity = traversable cells / 10, clamped between 2 and 30, and never above the number of parking places. Why density rather than parking places: ANADOLU has 144 loading cells, but putting 144 trucks into 266 traversable cells is not flow, it is gridlock. In multi-agent path planning, congestion rises sharply above roughly 10% agent density.
The same number sets two things at once: how many bookings may overlap, and how many route bands the RL state vector carries. That way every truck the scheduler allows is also visible to the policy. When capacity is full, a new booking moves to the first moment that frees up. Current values are in the table below.
The slot is reserved before the route is computed — because a route only means something inside the window it will run in. The order is:
t_end) and the schedule is
reflowed.Reservation times are stored as seconds from the start of the working day, and planning loads only those whose time window overlaps: a truck that finished at eight in the morning does not block an afternoon customer. The scope is the same day and the same warehouse.
ACO's cost grows quadratically with order size: at every step it tries all remaining products, and for each product all possible stopping points. A 1-product order is ~200 route searches; a 16-product order is ~27,000.
Bounded search (the default) holds the total number of searches fixed and reduces the round/ant counts if it has to. Full search performs the requested 10×10 exactly; on large orders it can exceed the time limit and be cut short. The estimate under the selector shows, live as the order changes, how many searches and how many minutes each option would take.
#12 and by colour.The video is produced as .webm (VP8).
VP9 was tried and was 4× slower for the same picture; these
frames are flat-shaded, so VP9's compression advantage buys nothing
here.
| Action space | How many distinct moves the model can choose from. It is map specific and decides which map a model can run on. |
|---|---|
| Episode | One job from start to finish: enter, collect the whole order, reach the exit. |
| Time-indexed A* | A pathfinding algorithm. Ordinary A* asks "which cells do I cross"; this adds "and when", so waiting and collisions can be accounted for. |
| Reservation | The record that says "this cell, in this time range, belongs to that truck". Collision avoidance is resolved from it. |
| Pheromone | The trail ants leave on good routes in ACO; it raises the chance those routes are chosen in later rounds. |
The values in the table below are not baked into this page, they are read from the running server — so the rule written here cannot drift from the rule the system enforces.