Gesture Control
    August 20, 2026

    How Gesture-Controlled Digital Signage Actually Works

    The interaction pipeline behind touchless signage, what the visitor physically does, and why the space in front of the screen decides whether it works — based on TapGest deployments in Prague and Sarajevo.

    How Gesture-Controlled Digital Signage Actually Works

    What gesture-controlled digital signage means in practice

    Gesture-controlled digital signage is a screen that responds to a person's hand movement in the air in front of it. A camera pointed at the space before the display is the only input device: it detects that someone has stepped up, follows their hand, and the application turns that movement into an on-screen action.

    That is a different arrangement from the three things it is usually confused with. A touchscreen needs the visitor to reach the glass, which means a physical queue at the surface and a surface to clean. A remote or controller needs an object that has to be handed out, charged and replaced. A phone-based interaction needs the visitor to scan something, open an app and look down at their own device instead of at the screen. Camera-driven gesture control removes all three dependencies — nothing is touched, worn, handed over or installed.

    This article is about how the interaction is put together and what it demands from the location. If you are looking at the commercial picture instead, that lives on our gesture control for digital signage page.

    The interaction pipeline, step by step

    Conceptually every one of our gesture installations runs the same loop.

    A visitor walks into the area in front of the display and is picked up by the camera. The camera image is processed to locate their hand. The application maps the position and movement of that hand onto a control — a spin, a swipe, a selection. The screen responds immediately, and the visual response is what tells the visitor that they, not a video loop, are driving the screen.

    The important property of this loop is that the feedback closes fast enough to be obvious. In the Samsung installation the scratch follows the hand in real time, so the connection between the movement and the screen is clear within the first second. If that first second does not land, people assume the screen is playing an advert and walk on — the loop is the whole product, not a feature of it.

    What the visitor physically does

    Across our deployments the physical action has stayed deliberately small, and always something a person will do in public without feeling awkward.

    At Fashion Arena Outlet in Prague the action is a spin: a shopper steps in front of the screen, raises a hand and sets a Wheel of Fortune in motion, then reads the result. Four moves, no registration, no promoter.

    For Jägermeister the same spin gesture drives a branded wheel whose prizes were exclusive merchandise, and for the Lasta Wheel of Treats at Bingo Shopping Center in Sarajevo the segments were Lasta's actual products — boxes of chocolates, bags of candy, baked goods — so people knew what was at stake before they played.

    The Samsung campaign in Sarajevo uses a different mechanic: the shopper swipes a hand across the air in front of the display and scratches away a covering layer to reveal a new smartphone underneath. Swiping continues until the device is fully exposed, and then the shopper can simply walk on — no form, no registration, no staff step to finish.

    What those four have in common is that the gesture vocabulary is one item long. Nobody has to learn a set of commands; there is exactly one thing to do, and doing it produces an immediate result.

    The space in front of the screen decides more than the software

    The camera defines an interaction zone, and a visitor only becomes a player once they are standing inside it. That single fact drives most of the site decisions.

    Screen placement is therefore an interaction-design question, not a facilities question. In a retail corridor the position of the display determines whether people naturally step into the zone or walk past just outside it. It also determines whether there is room to raise an arm without stepping into passing traffic — an interaction that requires a shopper to block a walkway will be abandoned regardless of how well it tracks.

    Unattended operation is the second constraint. These screens usually stand in an open corridor with nobody next to them, so the on-screen cue has to do the explaining. The wheel that is already turning when you approach is a cue; a static "raise your hand" caption is much weaker.

    A useful side effect of camera-driven interaction is the audience it creates. Because the visitor stands back from the glass rather than at it, one person plays while the next ones watch from a step behind and learn the mechanic before their turn. In a mall corridor that visible queue is part of why the format works — and it only exists because nobody is pressed against the screen.

    Touchless and touchscreen are not the same product without glass

    The obvious differences are hygiene and maintenance: a touchscreen in an open corridor has to be cleaned, and it adds a surface the operator maintains. Camera-driven signage leaves the display as a display, because nothing on it is handled by the public.

    The less obvious difference is the shape of the interaction. Touch supports precise, dense interfaces — lists, keyboards, small targets. Gesture does not, and pretending otherwise is the fastest way to a bad installation. Gesture is good at one coarse, visible action performed from a couple of steps back, in front of an audience. That is why our gesture work is wheels, scratch reveals and single-action games rather than menu-driven catalogues. The full comparison deserves its own article; for now, treat the distinction as coarse-and-public versus precise-and-private.

    What we have seen go wrong

    These are implementation issues we design around, not measured findings.

    Too many gestures. Every additional command needs to be taught on screen to someone who gave the display two seconds. One action is a game; three actions is a manual.

    A screen placed where nobody can stand. The interaction zone has to overlap with a place where a person can comfortably stop, which is not always the wall the operator wanted to use.

    Rounds that are too long. A single spin resolves in seconds by design, which keeps a corridor queue moving. An experience that occupies one visitor for minutes converts the people behind them into passers-by.

    Treating the campaign as code. Segment artwork, prize weighting and result screens are configuration in our builds, which is how the same mechanic carried Jägermeister merchandise in one centre and Lasta's product range in another. If a new campaign requires a rebuild, the format stops being economic.

    Hiding the reward. In Sarajevo the Lasta wheel showed the real products rather than abstract prize tiers, so people understood what they were playing for before they committed to being watched in public.

    Where the format fits

    In our experience it fits high-traffic public space with walk-past audiences: shopping malls and outlets most of all, and product launches where the reveal itself is the reason to stop, as in the Samsung case. It fits less well where the visitor needs to read, compare or enter data — that is touchscreen or phone territory.

    If the goal is a photo rather than a game, the relevant format is AR digital signage, which we cover in our tour of real AR deployment formats.

    Frequently asked questions

    Does gesture-controlled digital signage require touching the screen?

    No. The interaction is camera-based: the visitor's hand movement in the air in front of the display drives the content, and the screen surface is never touched.

    Does it need a special controller or wearable?

    No. In our deployments hand tracking is derived from the camera image alone — there is no controller, wearable or additional sensor to hand out or calibrate.

    Can gesture interaction run on an existing digital display?

    Yes, that is the usual case. At Fashion Arena, Samsung Sarajevo and Bingo Shopping Center the application ran on the venue's own signage with a camera facing the area in front of it.

    What kinds of interaction work well with hand movement?

    Single coarse actions: spinning a wheel, scratching or swiping away a layer to reveal content, triggering a game round. Dense interfaces with small targets are better suited to touch.

    Where is gesture-controlled signage most useful?

    In open public space with walk-past traffic and no permanent staff at the screen — mall and outlet corridors, in-centre promotions and product launches.

    Want to Learn More?

    Discover how we can create innovative AR experiences for your brand