Most of Meta's infrastructure runs on a question that sounds trivial: which thing goes where? Which shard lives on which server, which task runs in which region, which traffic flows through which path. On October 6, 2026, Meta's engineering team announced it is open-sourcing Rebalancer, the library it has used for more than nine years to answer those questions at scale.
This is not an AI model release, and we are including it because the problem it solves, placing work onto limited hardware under constraints, is exactly what AI infrastructure teams struggle with as clusters grow. This explainer covers what an assignment problem is, how Rebalancer is built, what the numbers mean, and where an ordinary engineer might use it.
TL;DR: the questions people are asking
| Question | Short answer |
|---|---|
| What is it? | A library that assigns objects to bins to optimize objectives under constraints. |
| License? | Apache 2.0. |
| Where? | github.com/facebook/rebalancer; Python package on PyPI. |
| Scale at Meta? | About 40M solves a day, 30+ problem formulations. |
| Speed? | P99 of 12 seconds at 265k objects and 3.2k bins. |
| Solvers? | Optimal MIP (Xpress, Gurobi, HiGHS) and parallel local search. |
| Debug tool? | Explorer, a Dockerized web UI. |
| Is it an AI tool? | No, but it fits GPU and workload placement problems. |
What an assignment problem is
An assignment problem asks you to place objects into bins so that all constraints hold and some objective is as good as possible. Rebalancer's vocabulary is intuitive:
- Objects and bins. Items to place and the places that hold them: tasks and servers, racks and data centers, shards and machines.
- Dimensions. Real-world attributes such as CPU, memory and storage, which objects consume and bins supply.
- Partitions and scopes. Groupings of objects and bins, so a rule can say "no more than two replicas of a shard in the same rack."
- Constraints. Rules that must hold, such as capacity limits.
- Objectives. Goals to optimize, such as balancing utilization or minimizing traffic across regions.
- Utilization. How much of a bin's capacity the objects assigned to it use.
The same shape appears outside data centers: assigning exams to rooms, nurses to shifts, tickets to agents, desks to employees. Meta itself lists meeting-room assignment, support-ticket routing and desk placement among Rebalancer's internal uses, alongside shard management, regional resource allocation, traffic routing, serverless function grouping and ML training workload balancing. One reply on X asked whether it could solve university exam timetabling. In principle the structure fits, though real timetabling has many special constraints, so you would need to model them.
How Rebalancer is built
Meta's central design idea is to separate how a problem is specified from how it is stored, solved and debugged. The specification is a three-level language:
- Modeling constructs such as objects, bins, dimensions and scopes.
- An expression API with operations like SUM, MAX and SQUARE, to build objectives and constraints.
- High-level specs, predefined recipes for common patterns, so most users do not write expressions from scratch.
Because the specification is abstract, the same problem can be solved by different engines, which brings us to the two solvers.
The optimal solver
This converts the expression graph into a mixed-integer program (MIP) and hands it to a commercial or open solver: FICO Xpress, Gurobi or HiGHS. It finds provably optimal answers, but its worst-case size grows with objects times bins, so it gets expensive on large problems.
The local-search solver
This works directly on the expression graph and improves an assignment by trying moves, such as moving an object to another bin or swapping two. Its neighborhood is small, on the order of objects plus bins, and it is designed to parallelize, evaluating millions of candidate moves per second. It gives up the optimality guarantee for speed and scale.
Having both is the practical choice. For small or critical problems you want the optimum. For a million objects that must be placed in minutes, you want a good answer quickly.
The numbers, and how to read them
| Metric | Value | Note |
|---|---|---|
| Solves per day | About 40 million | At Meta |
| Problem formulations | 30+ | Distinct uses |
| P99 solve time | 12 seconds | 265,000 objects, 3,200 bins |
| Large-problem average | 171 seconds | Over 1 million objects, 5,000 bins |
| Time in production | 9+ years | Per Meta |
A common reaction in the replies was skepticism that "12 seconds with 300k objects" solves anything useful. It does, because at hyperscale this is the work: placing hundreds of thousands of shards or tasks, then re-placing them as machines fail and load shifts, with solves running continuously. A 12-second P99 means almost every re-plan finishes quickly enough to act on. Another reply noted that nine years in production is itself valuable, because a solver that survived many surprise constraints is more trustworthy than one that only looks good on benchmarks. That is a fair point, though it also means the library reflects Meta's needs, which may not match yours.
Explorer: debugging the "why"
One underrated piece is Explorer, a Dockerized web interface for inspecting a problem and its solution. The question people ask first when an optimizer's output looks odd is "why did this job land on that server?" Explorer is meant to answer it by showing constraints, which ones were binding and how the decision was made. Optimization tools fail in practice when nobody can explain their output, so this is a meaningful feature, not decoration.
Where it is useful outside Meta
You do not need 40 million solves a day. Smaller teams could use it for:
- Placing replicas or shards across servers and zones with spread and capacity rules.
- Balancing ML training or inference workloads across GPUs by memory and compute, one of Meta's listed use cases. If you run several models on a cluster, deciding which GPU hosts which replica is an assignment problem.
- Scheduling batch jobs on a fixed pool of machines.
- Routing support tickets or tasks to people by skill and load.
- Allocating capacity across regions to cut latency or cost.
If you run Kubernetes, note that its scheduler places pods one at a time by heuristics, while a global solver considers many placements together. A reply compared Rebalancer to Google's Slicer, an auto-sharding system, and to Databricks' open-sourced Dicer, which are related in spirit. Rebalancer is a general solver library, not a full auto-sharding service, so it complements those rather than replacing them.
For context on running AI workloads at scale, see our posts on open-weight AI and the Kubernetes moment and multi-agent orchestration patterns, which face placement and scheduling choices of their own.
How to try it
A sensible first experiment, using Meta's published materials:
- Install the Python package from PyPI, or clone
github.com/facebook/rebalancer. - Start with a small, familiar problem: ten tasks and four machines with memory limits.
- Use a high-level spec for the pattern closest to yours, instead of writing expressions first.
- Solve with the optimal solver using HiGHS, which is open source, and note the result and time.
- Scale up to thousands of objects and compare with the local-search solver. Watch quality versus time.
- Open Explorer to see which constraints bind and verify the result makes sense.
- Add a realistic constraint, such as spreading replicas across racks, and observe how the solution changes.
- Compare against your current approach, whether a script, a heuristic or a scheduler's default, on utilization and fairness.
If you plan to use commercial MIP solvers like Gurobi or Xpress, remember they have their own licenses and costs. HiGHS is the free option.
Limits and honest caveats
- It is a general-purpose library. You still must model your problem correctly. A wrong constraint produces confidently wrong placements.
- Optimality at scale is not free. The exact solver can be slow on very large problems, and local search may settle for a good but not best answer.
- Meta-shaped assumptions. Nine years of production use shapes the API around Meta's problems, so some patterns may feel unfamiliar elsewhere.
- Adoption cost. Introducing a global optimizer into an existing scheduling system is an engineering project, not a drop-in swap.
- We have not run it. Everything here comes from Meta's announcement and engineering post. Check the repository for current documentation and limits.
What this means for what you build or pay
For most application developers, nothing changes. For anyone operating clusters, especially GPU fleets where idle capacity is expensive, a proven open-source solver lowers the cost of replacing hand-written placement heuristics with something measurable. A small improvement in utilization at scale is worth real money, and the ability to explain placements helps with trust. Treat it as infrastructure to evaluate, not as a headline.
Related reading
- Open-weight AI and the Kubernetes moment
- Multi-agent orchestration patterns
- Hugging Face RL environments and OpenEnv
- MacBook vs dedicated GPU for local LLMs
- Data centers: water, electricity and real impact
- Agent teams and the permission bottleneck
Primary: Engineering at Meta, "Rebalancer: a generic, high-performance library for assignment problems" (September 21, 2026) · the facebook/rebalancer repository · the companion OSDI 2024 paper on hyperscale resource allocation
Details are accurate as of October 7, 2026 and come from Meta's engineering blog and announcement. We have not run the library. Check the repository for current licensing, documentation and supported solvers.
