Known Issues & Limitations
What does not work yet and what works with limits. Use it to plan around gaps before you deploy. Each entry describes the current code and the current published artifacts.
Platform
- No Kubernetes support. The supported deployment targets are Docker Compose and HashiCorp Nomad;
pfy platform render --targetaccepts onlynomadorcompose. pfy platform deploysupports--target nomadonly. For Compose, runpfy platform render --target composeand thendocker compose up.- Single-node Nomad focus. The job files and guides target a simple, single-node Nomad setup. Multi-region and federated clusters are untested.
Published images
The optimizer image requires a ghcr login
The autoscaler release workflow publishes the optimizer as ghcr.io/productifyfw/autoscaler/optimizer:<version> and :latest, and that package is private. An anonymous docker pull fails: the registry refuses to issue an anonymous pull token for repository:productifyfw/autoscaler/optimizer:pull, and the tag list answers 403. ghcr.io/productifyfw/autoscaler/nomadscaler, which carries the per-commit and nightly tags of the strategy plugin, is private in the same way.
Until the packages are made public, log in first with a GitHub token that has read:packages and access to the ProductifyFW organization:
echo "$GITHUB_TOKEN" | docker login ghcr.io -u <github-user> --password-stdin
docker pull ghcr.io/productifyfw/autoscaler/optimizer:latestThe older ghcr.io/productifyfw/autoscaler-optimizer package is public but stops at 1.1.6, which predates the planning-horizon, cost-unit and task fixes; do not use it. The release-tagged strategy plugin, ghcr.io/productifyfw/autoscaler-nomadscaler (1.1.6, 1.1.7, 1.2.0, latest), is public.
Manager
pfy pilet upload cannot authenticate against a rendered Manager vhost
The Manager's pilet management routes (/pilets/:projectId[/:name/:version]) run the proxy-secret check, then the PAT middleware, then the identity-header middleware. The identity-header middleware takes the user only from the X-Token-Subject / X-Token-User-Email / X-Token-User-Name headers the proxy injects after a portal login and never consults the user the PAT middleware resolved. A valid PAT therefore passes the second middleware and is ignored by the third, which answers 401 no authenticated user.
The proxy does not rescue it. The middleware that turns a PAT into X-Token-* headers is productify_auth, which is per-application and fails closed without an application access decision, and the Manager's own vhost does not run it.
The Manager UI is unaffected: its same-origin fetches to /pilets/... carry the portal cookie, and the vhost's catch-all authorization plus header injection produces the headers the Manager wants. Giving /pilets/* a pre-portal route does not help and breaks the UI upload too, because a pre-portal handle skips the portal. Making pfy pilet upload work is a Manager-side change: the identity middleware has to accept the PAT context.
Workaround: upload through the Manager UI, or point the CLI at a Manager entrypoint that is not fronted by a rendered vhost.
The audit system role is inert
Every user carries a system role: normal, audit or administrator. administrator is enforced: it gates granting system roles, the audit export and evidence routes, the access-review routes, pfy proxy sync and the proxy-status dashboard panel, and it outranks project roles for operator resources. It does not make an administrator a member of any customer tenant; tenant-scoped permissions are never overridden by the system role.
audit is read by no code outside the generated enum. It grants nothing and restricts nothing. Use a project readonly role for observers instead.
Self-service join policies provision only into an unambiguous tenant
A project's join policy is invite_only, open or domain_allowlist. invite_only always works: every grant flows through an explicit invite. open and domain_allowlist admit an uninvited user only at OIDC login through an application vhost, and only when the target is unambiguous: the application's project must have exactly one tenant with that application enabled and exactly one tenant role flagged as the default member role (setProjectDefaultMemberRole). The user then receives that tenant role plus a readonly application grant, in one transaction. If there are several candidate tenants, or no default member role, nothing is granted and the login log records why; the project behaves as invite_only. On a multi-tenant project the two policies are therefore a no-op. Project (operator) membership is never granted by a login.
CLI
Structured output is for list and get commands only
The root --output/-o flag selects text (default), json or yaml, and under -o json|yaml diagnostics go to stderr so pfy -o json project list | jq sees clean JSON on stdout. Mutating commands such as create print human-readable text only and ignore -o. A few subcommands define their own flag: pfy deployment get -o yaml|json picks the spec format, and pfy nomad render --out <file> names the output file. See Output formats.
Integration libraries
Backend identity headers are trusted unless a proxy secret is set
@productifyfw/node-express and the Go library read X-User-* / X-Tenant-* unconditionally by default. Set proxySharedSecret (Node) or MiddlewareConfig.ProxySharedSecret (Go) to the proxy's secret: the middleware then requires a matching X-Proxy-Auth header (constant-time compare) before reading any identity header, answers 401 otherwise, and turns strict mode on unless strict is set explicitly. Without the secret, keep the application port reachable only from the proxy; anyone who can reach it directly can forge the identity headers.
Pilet signatures cover only the bundle bytes
@productifyfw/core's loadPilets verifies each bundle's SHA-256 against the feed item's hash and a detached Ed25519 signature (signature, key resolved by keyId) over the same bytes before evaluating anything, and the Manager signs the exact stored bytes at upload. The signature covers only the bytes, not the pilet's name, version, project or tenant. A genuinely signed older bundle can be served under a newer version (a downgrade), and a bundle signed for one project verifies in another, because the keyset is process-global. Closing this means signing a {name, version, project, sha256} manifest, a coordinated Manager and frontend change.
Two consequences of the current design to plan for: a feed of your own must emit hash, signature and keyId on every item or every pilet is refused, and a refusal is delivered to onError and nowhere else, so without an onError callback a tampered or unsigned bundle looks the same as "no pilets configured".
React has no comms binding
The notification bell and inbox ship only as the Vue-based @productifyfw/comms-pilet. @productifyfw/react exposes ProductifyProvider, the use* hooks, createReactPiletApi, loadReactPilets, PiletExtension and getPiletPage, and nothing for comms.
createI18nComposable lives in @productifyfw/vue. @productifyfw/core still exports a symbol of that name, but it is a deprecation shim whose body throws; import it from @productifyfw/vue.
window.__PRODUCTIFY__ carries auth_type, not authType
Every other top-level key the proxy injects is camelCase (tenantRole, tenantPermissions), but the authentication type is marshalled as snake_case auth_type; available_tenants and logout_url are also snake_case. authType is the intended canonical key and is not on the wire.
@productifyfw/core handles both: ProductifyFrontendConfig declares authType and a deprecated auth_type, and getAuthType() reads the canonical key first, then the legacy one. Call getAuthType(); if you read the raw global, read auth_type. The @productifyfw/dev shim and pfy dev proxy both emit auth_type, matching the proxy. The rename is a cross-repo change (proxy emits authType, core drops the fallback at 1.0) and both sides must move together.
Autoscaler
Scale-to-zero is not reachable end to end
min = 0 is documented in several places as the switch that turns scale-to-zero on. It is not sufficient on its own, and on the supported deployment path it cannot be set. Two different min = 0s exist, at two layers.
Layer 1 is the Manager's scale_to_zero deployment field. model.DeploymentScaling carries it, and it is the write-time gate that permits min: 0: a spec with min == 0 and the flag unset is rejected. The flag has no producer. DeploymentScalingInput exposes only enabled, min and max, the resolver carries the stored value across rather than accepting one, and pfy deployment set mirrors that. So no deployment can be given min: 0 through the API or the CLI, pfy nomad render (the only thing that emits a Nomad scaling stanza) never renders one, and the proxy renderer, which emits activator <env> inside productify_protect only when the primary deployment carries the flag, never emits an activator on a Manager-rendered vhost. This is deliberate; the schema comment on DeploymentScalingInput explains why.
Layer 2 is min = 0 in hand-written Nomad job HCL. The productify-scaler strategy accepts it, so a job written by hand, such as the reference fixture autoscaler/nomadscaler/config/test.hcl, does reach the optimizer with min_replicas = 0. The optimizer plans milp_startup_delay + ceil(cache_size / step_seconds) steps and returns the actionable tail, so at the shipped cache_size = 10 every response is a real decision. Driving POST /optimize with the shipped optimizer/config.ini and the shipped fallback allocation (fallback_cpu_alloc = 100 MHz, fallback_mem_alloc = 16777216 bytes) gives:
Scenario (min_replicas = 0) | cache_size = 10 | cache_size = 30 |
|---|---|---|
| idle app at 1 replica | [1] | [1, 1] |
| app at 0 replicas, demand far above capacity | [2] | [2, 4] |
| app at 0 replicas, only held cold-start requests | [1] | [1, 1] |
| app at 0 replicas, genuinely idle | [0] | [0, 0] |
The default cache_size scales up at every load. It does not necessarily scale down:
- One actionable step is one step of ramp. At
cache_size = 10the response is a single value, so a count moves by at mostmilp_max_scale_up/milp_max_scale_down(2 each by default) per evaluation. - Scaling down has to pay for itself. Releasing a replica costs
milp_shutdown_cost(0.3) once and savesreplica_coston every step of the actionable horizon, so a replica is released only when the saving over that horizon reaches0.3.replica_cost = cpu_alloc × cost_cpu_mhz_hour + (mem_alloc / 1024³) × cost_mem_gb_hour; at the shipped coefficients the memory term is negligible, so over one actionable step the threshold is 300 MHz of allocated CPU ([1]at 299 MHz,[0]at 300 MHz). At the shipped 100 MHz fallback an idle app holds its replica atcache_size = 10and30and first reaches[0, 0, 0]atcache_size = 45.cost_cpu_mhz_hourandmilp_shutdown_costare inherited values whose ratio is the single-step scale-down threshold; choosing them together is an open pricing decision recorded inautoscaler/optimizer/docs/cost-calibration.md.
What brings a zero-replica app back is not the optimizer. The proxy activator engages on a transport-level upstream failure when cached Manager state says the deployment is scale-to-zero eligible with no healthy backends, calls POST /api/proxy/wake, and the Manager scales the Nomad job to 1 directly. That path bypasses the optimizer, which is what keeps HTTP cold starts working despite the APM blocker below. It depends on the activator being rendered (Layer 1) and on PFY_ORCHESTRATOR_KIND=nomad; the noop default answers wakes with 202 and scales nothing. See Configuration.
The proxy gauge pfy_activator_pending_requests, the count of requests held through a cold start, is one of the range queries the optimizer issues and is consumed as an additive interactive demand term weighted 1.0. It has to be additive because the request-rate term is rate × avg_response_time × W and the response-time histogram is only observed when a handler returns, so it reads 0 for as long as requests are held. It does not wake a sleeping app, and on the supported deployment path the activator that populates it is never rendered. What it prevents is the next evaluation, seconds later and mid-drain, reading the app as idle and scaling it straight back to zero.
Still open: the sticky floor a wake raises so the optimizer will not scale the app back to zero mid-drain has no reader. It is an in-process Go store in the Manager (PFY_ORCHESTRATOR_PENDING_FLOOR_TTL sizes it) whose intended consumer is the Python optimizer in another repository, and no transport carries it there. Nothing but the policy cooldown holds a woken deployment above zero.
A zero-replica app is never evaluated: the APM errors before the strategy runs
Upstream of everything the optimizer decides, the shipped scaling policies never reach the productify-scaler strategy while an app sits at zero replicas.
Both shipped policies, the reference fixture autoscaler/nomadscaler/config/test.hcl and every job pfy nomad render emits, set source = "nomad-apm" with query = "avg_cpu-allocated". In nomad-autoscaler v0.5.0 (the plugin's pinned dependency) that query walks the job's allocations and skips every allocation that is not running; with no running allocation the sample list is empty and the APM returns metric not found. The check runner returns failed to query source before the strategy plugin is called, and the policy handler logs the error and moves on. The strategy's own "APM returned no metrics; skipping evaluation" guard never fires because the strategy is never reached. The target check does not stop it: a job at count = 0 is still Ready (the target reports Ready: !JobStopped), so the handler proceeds to the APM query and fails there.
Consequences for scale-from-zero:
- HTTP traffic survives, because the wake does not come from the autoscaler: the proxy activator calls the Manager's
POST /api/proxy/wakeand the Manager scales the job to 1 directly. See Scale-to-zero is not reachable end to end. - Queued and scheduled work does not. An app at zero whose only demand is a trigger, a CRON job or a queue backlog has nothing driving an activator, so nothing wakes it and nothing evaluates it. No supported policy
sourcereads thepfy_*series directly, so this cannot be worked around from the job file.
A policy without task cannot measure its own capacity
The strategy block carries two identities. metric_app_name selects the pfy_* demand series by the app Prometheus label, which is the application UUID stamped by the proxy and the Manager's queue executor. group and task select the nomad_client_allocs_* series that measure per-replica capacity, where the Nomad client exports the job's own task group and task names. task is optional on the wire, and a policy that omits it makes the optimizer fall back to metric_app_name as the task selector. That matches a real series only where the Nomad task happens to be named after the app label, which is never the case for a pfy-rendered job, where metric_app_name is a UUID. app_aliases cannot mask it; the allocation query does not go through the alias path.
What an affected app does depends on its replica count:
| Replicas | Behaviour |
|---|---|
running (> 0) | 503 {"error": "metrics_unavailable"} on every evaluation. The optimization is lost loudly and the plugin falls back to its own heuristic; the count is not sized by the optimizer at all |
| zero | The one silent case. Capacity comes from the last-known allocation for the (group, task), else from fallback_cpu_alloc / fallback_mem_alloc (100 MHz / 16 MiB at the shipped defaults), so a scale-from-zero decision is sized from constants rather than a measurement. Counted by pfy_optimizer_alloc_fallback_used_total |
How to tell whether you are affected: no `task` configured for check in the agent log (once per check), or No `task` in the scaling policy for app=... in the optimizer log (once per app). Either means the policy has no task line and the fallback is in play.
Fix: set task = "<the Nomad task name>" in the strategy block. pfy nomad render from pfy 2.0.0 emits task (and group) next to metric_app_name; a job rendered by an older pfy has no task line and must be re-rendered and redeployed, or hand-edited. The key is honoured only when both the strategy plugin and the optimizer are at 1.2.0 or later: an older plugin does not forward it, and an older optimizer ignores it. Pointing metric_app_name back at the task name is not a workaround; it breaks the demand query instead.
If you run into an issue not listed here, please open an issue on GitHub.