App development

How an app actually works

An app is a program installed on the device itself, whose interface and logic run on that device, and which talks to a server for every piece of fresh data. Most of what a user calls a slow app happens inside that conversation with the server, not in the code sitting on the phone.

  • Lesson 1 of 8
  • Beginner
  • Free, no signup

The five layers of an app, and the border every decision is taken on

  • User interfaceThe only layer the user sees, and the only one somebody tells you about when it breaks.
  • App logicDecides what each touch does and whether the data it needs is on the device or has to be asked for.
  • The API contractThe real border. Above it belongs to one device, below it is shared between all users.
  • The backendHome to the logic that must not sit on a user device: prices, stock, permissions, payment.
  • The databaseKeeps things. The user never talks to it directly and must not be able to.

This layering is a learning order, not a standard, and the names differ from team to team. The top two layers are on the phone and the bottom three are on the server.

Last checked: Facts and tool names in this lesson are re-checked against their sources on this date.

What happens when you tap a button?

One tap sets off four things in a row. First the interface code works out where the screen was touched and which button that touch belongs to. Then the app logic decides what data this button needs and whether it is already on the device. If it is not, an HTTPS request goes to a server. Finally, whatever comes back is turned into a screen.

Three of those four finish in a few thousandths of a second. The third one does not. The round trip to the server is almost always the longest part, and if you measure it once yourself you never forget it again. This command sends one request to this very site API and prints its stages separately:

curl -s -o /dev/null -w 'dns=%{time_namelookup} tls=%{time_appconnect} first_byte=%{time_starttransfer} bytes=%{size_download}\n' "https://rgb.ir/wp-json/wp/v2/posts?per_page=1&_fields=id,title,link,date"

We ran it three times in a row. Resolving the domain name took between 1.4 and 4.2 thousandths of a second, the secure TLS handshake finished by 31 to 49 thousandths, and the first byte of the answer arrived between 167 and 194 thousandths. Everything that came back was 353 bytes. A few hours later we took the same command again and this time the first byte came in between 180 and 251 thousandths, while the size of the answer stayed at exactly 353 bytes. The number worth remembering is the 353, not the times.

Look at the proportions. Setting up the secure connection took roughly a quarter of the time and the rest is how long the server needed to build the answer. Through all of it the app did nothing at all except wait. This is the exact place where a user says the program is slow and a developer goes looking for interface optimisations.

A phone with three glass layers rising in red, green and blue and a white beam to a distant server

The path of one tap, and where the time goes

  1. 1

    The screen is touched

    The interface works out where the touch landed and which button it belongs to.

  2. 2

    The logic decides

    Is the data already on the device? If it is, the job ends right here.

  3. 3

    The HTTPS request

    This is where the time genuinely goes. The secure handshake is part of this step too.

  4. 4

    The server builds the answer

    API answers are not cached, so every tap is a full execution on the server.

  5. 5

    The screen is redrawn

    The lighter the answer, the shorter this step, even on a cheap phone.

On this very site the third step returned its first byte between 167 and 194 thousandths of a second. The other three together do not reach that, though on a weak phone the last step becomes visible too.

Which layers is an app built from?

Five layers, and the border that falls between two of them matters more than the layers themselves. The user interface is what gets seen and touched: buttons, lists, forms. The app logic decides what each touch should do and what to display. Both of those run on the phone itself.

The third layer is the API, the contract that says at which address and in what shape the app asks the server for data, and in what shape the server answers. On the far side of that contract the backend sits on a server running the logic that must not live on a user device: prices, stock, permissions, payment. Underneath it a database keeps things.

The real border is between the second and the third layer. Everything above it belongs to this one device only; everything below it is shared between all users. Every hard question that comes up in an app project turns out to be this single question: should this job happen above the border or below it. Get the answer wrong and a customer edits the price on their own phone.

The layering above is a learning order, not a standard, and the names differ from team to team. What does not differ is that border.

Native, web and hybrid are three different things

When a client says they want an app, they almost always mean an icon on the phone screen. Three entirely different technologies give you that icon, and their differences only show up when the network drops or when you need to ship a new version.

Native means a program compiled for that operating system and installed as a package. Android own documentation puts it precisely: Android apps are written in Kotlin, Java and C++, and the SDK tools compile the code and resources into an APK or an App Bundle; the APK is the file a device uses to install the app, and an App Bundle cannot be installed on a device at all because it is a publishing format rather than an installation one. That single sentence becomes useful later, at publishing time.

Web means the thing that opens in a browser. Nothing is installed, an update reaches everybody the moment it ships, and access to the phone hardware is limited. A PWA is a version of that which, in MDN definition, is built with web platform technologies but provides an experience like a platform-specific app: like a website it runs on many platforms from a single codebase, and like an installed app it can be installed on the device, work offline and in the background, and integrate with the device.

Hybrid means a web application packaged inside a native shell. From the outside it looks native and on the inside it is web. It works well for forms and content, and it shows its ceiling wherever there is heavy animation or continuous work with the camera and the sensors.

Our position is simple: choosing between these three is not a decision about technology, it is a decision about the situation your user is in when they use the program. Someone working in a warehouse with no signal and someone sitting in an office on wi-fi get two different answers.

The four places where native and web genuinely part ways

  • Installed or opened

    Native is a package that gets installed and makes an icon. Web just opens at an address, and that is the lowest friction of entry there is.

  • What happens with no connection

    Native can keep a copy of the data and show something. Plain web shows an error page, unless it is a PWA built for exactly this case.

  • Where a new version arrives from

    Web reaches everybody the same moment. Native has to pass through a store, and there are always users who stay on an old version.

  • Access to the hardware

    Native has full access to the camera, the sensors and background work. Web has part of that, and its boundary is set by the browser rather than by you.

The line that separates each pair

Hybrid deliberately has no place in these four cells, because in each one it picks a side: it installs like native and inside it is web.

What stays on the phone and what lives on the server?

Two rules finish the job. Anything two people must see identically stays on the server: prices, stock, orders, messages. Anything the user must be able to see with no connection needs a copy on the device: settings, drafts, the last list they looked at, and the sign-in marker that saves them typing a password every time.

That sign-in marker, called a token, is the sensitive spot. While it is on the device the user stays logged in, and that is the convenience. If the phone is lost the token is lost with it, so it has to be revocable from the server side. That is something only the backend can do and no amount of code on the phone replaces it.

There is an unpopular thing to say about offline work, too. Most projects that ask for an offline mode actually want the user not to see an error screen while they have no signal in a lift. What they need is a read cache, and that is cheap. Real offline means the user can change something without a connection, and then you have to decide who wins when two people change the same thing at once. That decision is expensive, and a project that does not need it should not buy it.

Why your app is exactly as fast as its API

An ordinary list screen shows ten rows. We asked this site own API for ten articles, once in its default shape and once saying we only need four fields:

curl -s -o /dev/null -w '%{size_download}\n' "https://rgb.ir/wp-json/wp/v2/posts?per_page=10"
curl -s -o /dev/null -w '%{size_download}\n' "https://rgb.ir/wp-json/wp/v2/posts?per_page=10&_fields=id,title,link,date"

The first answer was 403,152 bytes and the second was 3,780 bytes. For a single article the same pair is 37,341 bytes against 353 bytes. We took five samples in a row and the sizes stayed identical; the times did not, and one of those five runs took 5.3 seconds on this live server. That gap is a lesson in itself: size is predictable and time is not.

Nothing changed on the app side. The list screen became a hundred times lighter because the request said what it needed. The rest of those 403 kilobytes was the fully rendered text of the articles, which a list screen never shows.

The same measurement carried one more thing. The API responses came with x-flying-press-cache: MISS and cf-cache-status: DYNAMIC, meaning that unlike this site HTML pages none of these answers is cached and every tap is a full PHP execution. A server that answers a web visitor comfortably is doing a different kind of work under an app.

So the right order of work is this: before choosing a framework, look at what your API returns. The next lesson on this path takes up that very choice, native or cross platform, and it is reachable from the lesson list on the app development path page. If you want to see what the same contract looks like from the web side, the lesson on what an API is opens it from the ground up.

The same screen, before and after the request says what it wants

bytes, measured against this site own API on 8 September 2026
  • One article 37,341 bytes353 bytes
  • A list of ten articles 403,152 bytes3,780 bytes

    The list screen is where the difference is made, because it asks for ten times the data.

the default requestthe request asking for four fields

These numbers are size, not time. Size stayed exactly the same across five samples, while the time reached 5.3 seconds once; on mobile data it is the size that shows itself.

The fast path, with AI

The fastest thing a model can do for you here is not writing code. It is pulling the data contract of the app out before a single line exists: what each screen shows, which address that comes from, and exactly which fields are needed. The same move that made the list screen a hundred times lighter in the last section.

  1. Write down the screens of the app by name, the way you would describe them to a person: sign in, order list, order detail, profile. No technology, no framework.
  2. For each screen write only what the user actually sees on it. If the list screen shows a title and a date, write those two and nothing more.
  3. Hand that list to a strong model with the recipe below and have it build the API contract. Pick a judgment class model rather than the fastest and cheapest one, because the work here is deciding, not translating.
  4. The last step is the one nobody does: ask the model to delete every field no screen consumes, and to say which fields it guessed. That single line turns the list from something that looks good into something that is light.

Copy-ready recipe

Role: API architect for a mobile app.

The screens of the app and what each one shows:
{screen 1}: {the fields the user sees on it}
{screen 2}: {the fields the user sees on it}
{screen 3}: {the fields the user sees on it}

What to do, in this order:
1. For each screen write one request: method, address, parameters and a sample JSON response.
2. In the sample response keep only the fields that screen displays. Delete every extra field.
3. Give a table of the fields you deleted and the screen that would have to bring each one back.
4. Separately, state which names or structures you guessed, and where I should verify them.
5. If data is shared between two screens, say which screen should cache it and for how long.

Write no number about speed or size. I will measure that myself.

Before you trust the output: The output of this recipe is a proposal, not a document. The model invents field names and even addresses, so every line of it has to be compared with what the backend actually returns; one curl against the real address finishes that job. And if the project has no backend yet, this contract is the input to building one rather than a substitute for it.

AI in this kind of work

On this topic models do two things genuinely well and one thing badly. Well: reading a long platform document, and pulling a data contract out of a screen description. Badly: any sentence that touches a number or a store policy. Our position is to put the model on the contract design side and keep it off the claim making side.

Tools that actually help

  • Claude Suited to exactly what the recipe above asks for: turning a description of screens into a data contract and, more importantly, deleting what no screen needs. Iran is on neither of Anthropic two supported-countries lists; we read that on their own page rather than measuring it.
  • Gemini Good at reading the long platform documents, and this lesson itself takes several of its sentences from the official Android documentation. Google own page says the Gemini web app works in more than two hundred and thirty countries and territories, and Iran is not on that list.
  • NotebookLM When the question is what an SDK document exactly says, this tool is more correct than an ordinary chat, because it answers only from the source you gave it. Google own help says it works in the same regions as the Gemini app, and Iran is not on that list.
  • GitHub Copilot It works inside the editor, and for an Iranian reader its access story is the most unexpected: GitHub own trade-controls page says a US Treasury licence covers its cloud services for developers resident in Iran, free and paid. We quote that and make no claim about payment.

Where it backfires

The first risk here is specific to apps and directly connected to the model. An app goes onto the user phone, so anything inside the package is in the hands of whoever holds that phone. Models routinely write example code with the API key sitting in the middle of it, because that is the only way an example runs; and then that example gets copied and shipped to a store. The OWASP Mobile Top 10 in its 2024 release puts Improper Credential Usage in first place and Insufficient Binary Protections in seventh. The rule that falls out of this lesson is simple: a key that has to stay secret has no place above the border and belongs in the backend.
The second risk is more ordinary. A model writing about mobile SDKs or store rules will invent a method name or a publishing condition in the same confident tone. Every sentence of that kind has to be checked against the vendor own documentation. How to pay for these tools from Iran is in the buying guide.

Sources: OWASP Mobile Top 10 (2024) Anthropic: supported countries GitHub and Trade Controls Google: where the Gemini web app is available

Where this advice stops

This lesson explains the common shape: an app with screens that talks to an API over HTTP. Games are not here, and neither are apps whose value is built on the device itself, such as camera image processing or a model running on the phone; in those the border this lesson stands on sits somewhere else entirely. One note about us is also needed: the depth of our first-hand experience is the server side of that border, the hosting and the API, and every number on this page is a server number rather than a number from a published mobile app.

From our own work

On this same server we keep a small must-use plugin called rgb-rest-raw.php that exists for one reason only. The security plugin Really Simple SSL opens an output buffer and rewrites every https://rgb.ir into https://rgb.ir across the whole response body. For HTML that is correct. On a JSON answer it is a silent corruption: our redirect checking tool exists to show that an http address redirects to https, and that rewrite made the first link of the chain read https to https, so the very thing the tool was built to find became invisible. The value stored in the database was right; only the output body was edited. No browser would have noticed. A program reading an address byte by byte does. The second item is of another kind: our own public API sits behind a captcha, and a plain request to /wp-json/rgb/v1/domain returned code 422 today with a message asking to wait for the security check to complete. On a website form that is correct, and it is exactly why an app cannot consume a website API as it stands and needs an authenticated route of its own.

Real follow-up questions

Do I definitely need a server to build an app?

No, if everything stays on the device itself: a calculator, offline notes, a measuring tool. The moment two devices must see the same thing identically, or a user has to sign in to an account, a server becomes necessary rather than optional. The rule in the fourth section marks that border.

What is the difference between an app and a mobile site?

Four things: an app installs and makes an icon, it can show something with no connection, it has full access to the camera, the sensors and background work, and a new version of it has to pass through a store. A site has none of that, and in exchange it has no friction of entry at all and its update reaches everybody the same moment.

Why is my app fast on wi-fi and slow on mobile data?

Because it is almost always a size problem rather than a code problem. Wi-fi hides a big payload and mobile data does not. The measurement in the last section shows exactly that: a ten row list screen from this site was 403,152 bytes in its default shape and became 3,780 bytes once the request named the four fields it needed, with nothing changed on the app side.