Sites and apps
How to deploy a Gradio app without Hugging Face Spaces
In short
You can put a Gradio app online on any host that keeps Python running: you need a requirements.txt with gradio and a server bound to 0.0.0.0 instead of 127.0.0.1. A share=True link does not work for this, because it is temporary and routed through Gradio servers. On Netrun the built-in build recognizes Gradio and starts it on the right address and port, so usually there is nothing to configure. Netrun has no GPUs, so large models are better called through an API while the server runs light ones.
- By default Gradio listens only on 127.0.0.1 and port 7860, so the app is not reachable from outside.
- The Gradio address and port come from the GRADIO_SERVER_NAME and GRADIO_SERVER_PORT environment variables, and explicit launch() arguments override them.
- A public share=True link is routed through Gradio servers and expires after one week.
- By default Gradio processes one call to a function at a time, and the rest wait in a queue.
- New Gradio apps on Hugging Face Spaces require a paid plan, except for two ZeroGPU apps allowed on free personal accounts.
People love Gradio because a model interface takes a dozen lines: an input, a button, a result. Then you want to show the demo to a client, a teacher or your followers, and it turns out that 127.0.0.1:7860 opens for nobody but you, while a share=True link stops working after a week. The usual answer is Hugging Face Spaces, but new Gradio apps there now require a paid plan, apart from a limited free option on ZeroGPU.
The good news is that Gradio is an ordinary Python web app, and it runs on any host that keeps Python running. Below is how to prepare the project, what to do with the address and port, and how not to run out of memory. If you have published a dashboard before, much of this will feel familiar from the Streamlit guide.
| Criterion | Hugging Face Spaces | Netrun | Your own server |
|---|---|---|---|
| How code gets there | Through the Space git repository, rebuilt on every commit | As an archive, a folder, from GitHub or from an AI editor | However you set it up |
| GPU | Paid GPUs and ZeroGPU with limits | None, CPU only | If you rent a machine with a GPU |
| Running Gradio | The platform is built around Gradio | The build detects Gradio and starts it on the right address and port | By hand: address, port, autostart, proxy, certificate |
| Sleep without visitors | Free hardware sleeps after about 48 hours without use | Sleeps on the free plan and wakes when opened, never sleeps on Pro | Never sleeps |
| File persistence | Space disk is not persistent by default | The /data folder survives code updates and restarts | Everything is on your disk |
| Best for | Public demos in the Hugging Face community and models that need a GPU | CPU demos and apps that call models through an API | Full control and your own GPUs |
List your dependencies in requirements.txt#
Next to the main file, usually app.py, put a requirements.txt with your libraries, including gradio itself. The host installs dependencies from it, and Netrun also uses it to recognize a Gradio app. If you use torch and the server has no GPU, install the CPU build: it is much smaller and installs faster.
Do not hardcode the address and port in launch()#
Gradio reads the address and port from the GRADIO_SERVER_NAME and GRADIO_SERVER_PORT environment variables, but explicit server_name and server_port arguments in launch() override them. On Netrun the built-in build starts Gradio on 0.0.0.0 and port 8080, so a plain demo.launch() with no arguments is enough. If your code has server_name="127.0.0.1" or server_port=7860, remove them, otherwise the link will not open. In your own Dockerfile, set server_name="0.0.0.0" and take the port from the PORT variable.
Remove share=True#
The share=True option creates a temporary public link through Gradio servers, and it stops working after a week. On a host you do not need it, because the app already has its own HTTPS address. Besides, such a link routes visitors to your app through third-party servers instead of directly.
Estimate how much memory the model needs#
A model is usually loaded into RAM in full, and large models easily exceed the plan limit, in which case the system stops the process and the logs show Killed. Netrun has no GPUs and everything runs on the CPU, so language models and image generation will be slow or will not fit at all. Light models such as classifiers and small text models work fine, while heavy ones are better called through the API of a service that runs them.
Keep downloaded models in a persistent folder#
If the app downloads a model at startup, without a persistent folder it downloads it again after every publish. Set the HF_HOME environment variable to /data/huggingface, which on Netrun is done in the Secrets tab. The Hugging Face library cache then lands in the persistent /data folder and survives code updates. Keep in mind that the model takes up space on your plan disk.
Publish and share the link#
Upload the project as an archive, a folder or from GitHub, put third-party API keys into Secrets and wait for the build. Netrun installs dependencies, starts the app and gives you an HTTPS link, and the live logs show model loading errors right away. If many people use the demo at once, keep the Gradio queue in mind: by default a function processes one call at a time.
Hugging Face Spaces is still a good home for public demos and models that need a GPU. If your demo runs on the CPU or calls models through an API, it is simpler to keep it on ordinary hosting: on Netrun Gradio starts from your code with no address or port setup, keys live in secrets, and models in /data are not downloaded again. On the free plan the app sleeps without visitors and wakes up when the link is opened, so the first visit after a pause takes a few seconds. For a demo you present live, Pro is the better fit. Try Netrun.
Common questions
Why does my Gradio app not open on the server?
Most often it listens on 127.0.0.1, which is the Gradio default, or launch() hardcodes port 7860 while the host expects another one. Remove server_name and server_port from launch(), or set the address to 0.0.0.0 and the port from the PORT variable. The exact cause usually shows up in the startup logs.
Can I just keep using share=True instead of hosting?
For a quick demo over a couple of days, yes, but the link through Gradio servers expires after a week and only works while your computer is running. A permanent address needs a host where the app runs without you.
Can a server without a GPU handle a large model?
Usually not: large language models and image generation are very slow on a CPU and often do not fit into the plan memory. A practical setup is to host the Gradio interface on the server and call the model itself through the API of a service that runs it. Light models work fine on a CPU.
Why does my app crash with Killed?
That is how the system stops a process that ran out of RAM, and it most often happens while loading a model. Use a smaller model, load it once at startup instead of on every request, or move the heavy part to an external API. If the app itself needs more memory, you need a plan with more resources.
How is publishing Gradio different from Streamlit?
Both are Python web apps and are published the same way: a requirements.txt, a main code file and a server bound to 0.0.0.0. Gradio is better for an interface to a single model with inputs and outputs, Streamlit for dashboards and reports. On Netrun both frameworks are detected automatically.
Can I run a Python project — Django, Flask or FastAPI?
Yes. Put your dependencies in requirements.txt — Netrun detects Python and builds the project for you. Django, Flask and FastAPI are supported; the app should listen on the port from the environment variable, and keys and database access are set as secrets rather than in the code.
Why does my website sometimes sleep?
On the free plan, websites sleep after being idle and wake up in a couple of seconds on the first request. On the Pro plan your project runs without sleeping.
Where do I set tokens and other secret values?
Every project has a Secrets tab where you set the values from your code — for example the token from BotFather. We store them encrypted: you can see the variable names, but the values are shown to no one, including you.