GenEye in a Box: Making Machine Vision Something You Can Just Ask For

Computer vision can help factories catch defects, check packaging and keep production flowing. Yet many small and mid-sized manufacturers still don’t use it. Technology exists but getting it running and turning an idea into a useful application often requires specialist knowledge such as selecting a model, setting up tracking, defining rules and connecting the results to a workflow. For a company with limited technical resources, those steps can make even a straightforward task difficult to implement.

GenEye is a research project that tackles this gap.
Our question is simple: what if setting up machine vision were as easy as telling a device what you want it to look for?

The problem we want to solve

The project focuses on manufacturing and food industry companies in South Ostrobothnia. Earlier development work in the region showed that these companies see the potential of computer vision for production efficiency and quality control. Adoption is held back by three things:

  • the need for specialized expertise,
  • competence gaps within companies,
  • the complexity of advanced technologies.

The idea: A vision system in a milk-carton-sized box

GenEye aims to build and pilot a computer vision solution that combines four ingredients in one compact, user-friendly device, roughly the size of a milk carton:

  1. Generative AI to interpret what the user wants.
  2. Algorithm generation so the system can produce the vision logic itself instead of requiring a specialist to hand-craft it.
  3. Voice control so operators can interact in plain language rather than through complex tooling.
  4. Edge computing so processing happens on the device, close to the production line.

The aim is a device that merges state-of-the-art technology with everyday usability. The people on the factory floor should be able to use it without becoming computer vision engineers.

From a request to a vision task

GenEye is being developed around a simple idea: users describe what they want to observe, and AI helps turn that request into a computer vision application. The proposed workflow connects spoken or written instructions with code generation, deployment, camera processing and understandable results.

The process begins with a user’s request. For example, an operator might say: “Count the bottles passing this point and notify me after every hundred.” A spoken instruction would first be converted into text using speech recognition. GenEye would then interpret the task, identifying the objects of interest, the conditions to follow and the results the user expects. When an instruction is unclear, the system should ask for clarification—for example, which point bottles should cross or how long an empty production line should remain idle before resetting the counter. Then, an AI coding agent would generate the code needed to perform the task.  Before deployment, the code would need to be checked and tested to confirm that it runs correctly and follows the intended rules. The checked application would then be deployed to the computing environment connected to the camera. Deployment would prepare the required software, models and settings so the application could receive and process the video stream. Once running, the application would analyze incoming frames and present its results to the user.

The interaction would continue beyond the initial request. A user could ask to change the notification threshold, select a different inspection area or stop the analysis. GenEye would interpret the change and update the configuration or regenerate and check the relevant code as needed. The aim is to make adapting a vision task a natural conversation, while keeping the resulting application testable and understandable.

Use cases

GenEye’s intended flexibility opens up several potential applications in manufacturing, logistics and workplace monitoring. The following examples illustrate tasks that could be explored, depending on the available camera views, suitable models and application-specific testing.

  • Object detection and identification. Locate products, packages, tools or other relevant objects in a camera view. An operator could ask the system to highlight bottles on a conveyor or identify whether the expected components are present at a workstation.
  • Object counting. Count items passing a selected point, entering an area or leaving a production line. Counting rules could include batch totals, notifications after a specified number of products and automatic resets between batches.
  • Object tracking. Follow products, carts or other moving objects across video frames. Tracking could help measure how long an item remains in an area, observe movement direction or identify an object that has stopped unexpectedly.
  • Packaging and assembly inspection. Check visible features such as the presence of a bottle cap, a label or an assembly component. The application could flag a possible missing or misplaced element for an operator to review.
  • Anomaly detection. Highlight departures from an expected appearance or activity pattern, such as damaged packaging, an unusual accumulation of products or an unexpected interruption in movement. These tasks require a clear definition or suitable examples of normal operation.
  • Workplace safety monitoring. Explore checks for visible protective equipment, entry into restricted areas or proximity between people and moving vehicles. Such applications would need careful validation and should support established safety procedures and human oversight.
  • Production flow monitoring. Observe queues, empty conveyor sections or recurring stoppages. Users could request notifications when products accumulate in a particular area or when no item passes a checkpoint for a defined period.
  • Warehouse and storage monitoring. Detect whether designated spaces are occupied, count visible packages or flag objects obstructing a marked route. These observations could help staff review changes and respond to situations requiring attention.

Stay tuned…

We are still developing the idea behind GenEye, and there is much more to come. We’ll share updates, prototypes and lessons learned soon.

GENEYE in a BOX – Seinäjoki becomes a pioneer in the business use of machine vision through voice control and generative artificial intelligence | Tampere University

Funding source

Etelä-Pohjanmaan liitto (EAKR), Seinäjoen kaupunki, Etelä-Pohjanmaan korkeakoulusäätiö

About the author

Afshin Dini

Project Researcher

Scroll to Top