Key Presets and Custom Policies
UpdateyesterdayUnverifiedAdded Oct 10, 2026
Keys now use permission presets instead of the ADMIN and API scopes. When you create a key, select one of seven presets: API , SERVERLESS , DEPLOY , COMPUTE , BILLING , READONLY , or FULL . If no preset matches your case, select the exact permissions the key needs. You can also limit a key to specific models and set it to expire.
Topics: Governance, Security, Pricing, Developer tools
More fal releases
Every fal releaseBilling Headers for WebSocket Endpoints
Shared WebSocket endpoints can now declare billing on their 101 upgrade response. Set x-fal-billable-units when the unit count is known at connect time. Set x-fal-billable-units-webhook to 1 when the count is only known at the end of the session. Then report the total with POST /requests/billable-units/{request_id} .
Switch Between App Endpoints in the Dashboard Playground
Deployed Serverless apps now include Testing > Playground in the app dashboard. Select an available endpoint, fill in its generated input form, and run a request without leaving the app. Multi-endpoint apps have one place to test their endpoints, using the existing Playground authentication and billing behavior. Open the guide
Platform MCP Server
fal's second MCP server , the Platform API MCP , is now available at https://api.fal.ai/v1/mcp/platform . Where the Run MCP lets your AI assistant build with models , this one lets it operate your fal account : connect Claude Code, Cursor, or any MCP-compatible client with your API key and ask "why did my last request to my-app fail?" , your assistant walks...
GPU Utilization in Runner Telemetry
Runner telemetry now includes a dedicated GPU Utilization chart alongside the existing CPU, memory, and VRAM charts, so you can see how much compute your GPUs are actually doing -- not just how much memory they hold.
Deploy Messages and Annotations
You can now attach a freeform message and custom annotations to each revision when deploying: Messages and annotations are set at deploy time and stay with the revision. View and search them on the app's Versions page in the dashboard, or fetch them via the revisions API .
App-Level Retry Configuration
Apps can now register a default per-condition retry budget at deploy time with the new retry_config deployment option, using the same format as the X-Fal-Retry-Config request header. Set it as a class attribute on fal.App or under the app's table in pyproject.toml .
New Serverless Usage Page
The new Serverless Usage page (Serverless → Usage) provides attributable, dimensioned reporting of serverless compute spend in the dashboard. Compute reported in machine-seconds , broken down per app, environment, and machine type. Spend is split by price category: Reserve, Burst, On-Demand, and List.
Redesigned Usage & Billing Page
The Usage & Billing page has been rebuilt to support usage attribution (the Settings "Overview" tab is now "Usage"). Filter by auth method, API key, and login username to attribute consumption to a specific credential or user. Normalized app names and revised charts improve readability and export.
Also shipped on Oct 9, 2026
fal in October 2026Google Workspace connector is available in Beta
The managed Google Workspace connector is now available in Beta in Databricks Lakeflow Connect. Use the connector to ingest audit activity from Google Workspace applications and services into Databricks. The connector uses OAuth user authorization and supports incremental ingestion.
Agentic Marketplace Discovery (Public preview)
Agentic Marketplace Discovery is now available in public preview. On the public Snowflake Marketplace Discover page in Snowsight, you can switch between Agentic discovery and Marketplace search . Describe what you need to CoCo to find and compare matching listings, or search for providers and listings directly.
All AI Gateway models with prepaid credits, backend features with your coding agent, and new ways to send feedback release - Oct 09, 2026
All AI Gateway models now available with prepaid credits. You don't need to request access to frontier models in the Neon AI Gateway anymore. If you're on a pai...
Use request tags in Unity Gateway service policies (Beta)
Custom service policies in Unity Gateway can now check caller-supplied request tags, such as requiring a tag before a model request or MCP tool call proceeds. See Request tags .
Use OpenJev as the evaluator for LLM-as-a-judge service policies (Beta)
You can now select OpenJev (Qwen3.5 4B) as the evaluator model service of a custom LLM-as-a-judge service policy. OpenJev classifies content without generating output tokens, so it typically responds faster than a chat evaluator. See Use OpenJev as the evaluator .
Code execution in Genie One (Public Preview)
Genie One can now run code in a secure, isolated sandbox to take on work that goes beyond SQL, such as more advanced analyses. Code runs on your behalf using your existing permissions, so Unity Catalog governance and data access controls continue to apply. See Code execution in Genie One .