Showing posts with label Prototypes. Show all posts
Showing posts with label Prototypes. Show all posts

Monday, August 9, 2010

Ultimate – A prototype search engine

Lately, I have been experimenting again with Axure on a prototype search engine called Ultimate. The engine is based on actual user requirements collected through a focus group study. I decided to prototype Ultimate, in order to perfect my skills in Axure. The tool enabled me to construct a high-fidelity and fully functional prototype within a few hours. Some of the features of the engine are:

  • It relies on full natural language processing to understand the user’s input.
  • The search algorithm is based on a complex network of software agents – automated software robots programmed to complete tasks (e.g., monitor the prices of 200 airlines, get ratings from tripadvisor.co.uk, etc)  - to deliver accurate and user-tailored results.
  • Some of the engine functionalities are discussed in the user journey shown below. The full functionalities are well documented, but for obvious reasons I can not discuss them in this post.

Screenshots:

 
User Journey:
 
 
 Research:

Run a usability testing of the above design against a more “conventional” search engine (e.g., Skyscanner ), and I am certain that the results will show the clear superiority of Ultimate. Of course, there is the need for a careful usability study in order to compare the designs but I am certain that Ultimate is superior in all usability metrics.

 

  

Friday, July 30, 2010

E-Learning Prototype

Below is the prototype of a e-learning system that I was asked to do by a company. As I can not draw, I decided to use Microsoft Word to communicate my ideas. There should be a good storyboarding tool out there that could help me to streamline the process.

The design below is based on existing and proven technologies that can be easily integrated into existing e-learning platforms. Codebaby, a company in Canada is already using avatars (such as those shown in my design) [1] in e-learning very successfully for several years. The picture in the last screen of the design is a virtual classroom [2] created in the popular Second Life platform.

 
Compare my solution with a “conventional” e-learning platform shown below. Although I do include several GUI (Graphical User Interface) elements in my work, it is obvious that : a) my interface is minimalistic with fewer elements on the screen. b) accessibility is greater, as instead of clicking on multiple links in order to accomplish tasks, you can simply “ask” the system using the most natural method you know - “natural language”. The benefits of avatar-assisted e-learning will become evident when the web progresses from its current form to Web 2.0 and ultimately to Web 3.0. For now, such solutions should at least be offered as an augmentation to “conventional” GUI-based interfaces. All companies want something more, for example something that would add easier access to module contents and  the “WOW” factor to their products. They just don’t know what it is until you show it to them.
 
1[2]
 
Although the proposed design is based on mature and well-tested technologies, I can understand if someone wants the purely GUI solutions. In fact, I would be more than happy to assist them. I have been working with GUI interfaces for several years, long before I developed an interest for avatar technologies. I developed my first e-learning tool back in 1998 (12 years ago). It was an educational CD-ROM about the robotic telescope platform of Bradford University.
 

[1] http://www.codebaby.com/showcase/

[2] http://horizonproject.wikispaces.com/Virtual+Worlds

Monday, July 26, 2010

MGUIDE Development Process

I thought it would be a good idea to try to explain the methodologies followed in the development of the MGUIDE prototypes. Having a focus mainly on the research outcomes, the development methodology followed was of little concern to the involved stakeholders. Trying to create interpersonal simulations like the ones found in real-life is a process mostly compatible with the a Scrum development methodology (shown below). I am planning to create a paper on the topic, and hence I will not say much in this post.

  800px-Scrum_process_svg

Source: Wikipedia

Gathering the requirements of the users can be done using a variety of ways. I followed a combine literature-user evaluation approach. One of my earliest prototypes was developed using guidelines found in the literature. The prototype was then evaluated with actual users and a set of new requirements was developed. These requirements are what the SCRUM refers to as the “product backlog”. Each spring (usually in my case 1-3 months) a set of the requirements were developed and tested, and then were replaced by a new set of requirements. Doing simulations of interpersonal scenarios gives you the freedom to augment the product backlog with new requirements quite easily. Using methods of research like direct observation and note taking, you can take notes on the interactions found in the scenarios that you want to simulate. My scenario was a guide agent and hence, I went to a number of tours where I made a number of interesting observations. Most of my findings were actually developed in the MGUIDE prototypes, but there are others that still remain in the “product backlog”. Of course these requirements and the work that was done in the MGUIDE is enough to inform Artificial intelligence models of behaviour in order to create completely automated systems.

This iterative process was then repeated prior the actual user research stage, where the full set-up of the MGUIDE evaluation stage was tested. I used a small group of people that tried to find bugs in the software, problems with the data gathering tools and others. The problems were normally corrected on site and the process was repeated again. Once I ensured that all my instruments were free of problems, the official evaluation stage of the prototypes started.

Closing this post, I must highlight the need for future research in gathering data about different situations where interpersonal scenarios occur. In reality different situations produce different reactions in people and this should be researched further. Only through detailed empirical experimentation we can ensure that future avatar-based systems will guarantee superior user experiences.

 

Friday, June 4, 2010

Rapid Prototyping, Which tool?

Below is an experiment to simulate one of the MGUIDE prototypes using Axure RP Pro. I wanted to see if Axure can be used in such applications, and how fast I could implement a working prototype.The dialogue window is clickable, as well as the buttons. All standard GUI (Graphical User Interface) elements of the MGUIDE interface can be quickly and easily implemented using this tool. The environment is drag drop, supporting adding interactions using visual elements (e.g., If then else statements with a few clicks). However, the absence of an internal scripting language means that I am limited by the number of interactional elements the program has already build in. I can’t build my own interactions, like for example the dialogue branching I easily implemented using VB.NET. Neither can I import my 3D avatar, for lets say a formal evaluation. In conclusion, Axure appears to be a tool to be used very early in the product cycle and mostly for low-fidelity prototypes. Later on, you you will need to build something complex in order to communicate your ideas more effectively. This is where I come in!  I can built highly complex prototypes, in the same amount of time using VB.NET/Adobe Director with just a few lines of code. However, if a a stakeholder is happy with this tool, I can be happy as well. The tool can be used virtually by anybody, let alone myself. 

System Preferences Wireframe

Prototype_page0

Main Screen Wireframe

Prototype_page1

Saturday, May 29, 2010

Prototype 5

This prototype is similar to the others, but it:

1) Features only brief descriptions about four locations of the castle of Monemvasia. This content has been carefully crafted to ensure that it is short (but not to short) in terms of length, analytical, and most importantly  simple to understand.

2) The Loquendo Kate text-to-speech voice speaks too fast even in low speed mode. To fix this, I have introduced a 1 sec delay between each sentence of the description. That makes the presentation to flow more naturally and the text easier to understand.

3) Features two types of virtual guides:

a) One virtual guide, that uses full gestures and happy facial expressions. This guide always looks at the user.

b) Another virtual guide that uses no gestures and her face appears always serious. This guide looks away from the user most of the times.

The screenshots below demonstrate both guides:

guides

Friday, May 28, 2010

Prototype 2

Prototype 2 (and prototype 1) were evaluated in Greece. The idea here was to build a system that offers information of variable complexity about the castle and most importantly allows the visitor to navigate freely in the castle.

A number of solutions were considered in order to allow free navigation:

a) Real – time landmark recognition. This feature is shown below (move to slider 10 1.58). The idea here, is to train the system to recognize a particular location (simply by taking a photo) and allow the user to initiate a presentation about the location by pointing the camera of the device to it. That would be extremely neat to have in the current prototype 2 implementation but the technology is probably still very expensive to purchase.  

b) QR-CODE technology. This feature is shown below (Greek language only). The idea here, is to tag each location on the castle with a QR-CODE (similar to the product bar code but it can hold textual as well as numerical information) and allow the user to initiate a presentation about a location simply by photographing the QR-CODE of the location. A more advanced version of the technology, allows a similar functionality to the real-time landmark recognition (i.e.,  real-time video QR-CODE recognition). It would be interesting to come up with a study to compare the performance and usability of both technologies. At the moment apart from cost (QR-CODE is far cheaper) I see no core differences between the two technologies.

Monday, May 24, 2010

Videos – Prototype 4 (2nd Demo)

    The video below demonstrates the two processing layers I constructed upon the VPF for better language understanding. An early development of the algorithm can be found here. Please note that the few minutes of delay at the beginning of the video are because of the initialization of the tagger.

    This system uses shallow parsing and deep syntactic processing to match the user’s input with the database. In particular, the following steps are taken to find the database phrase closest to the input:

    Stage 1: Shallow Parsing

  • Replace contractions

[didn't, 'll, 're, lets, let's, 've, 'm, won't, 'd, 's, n't] with [did not, will, are, let us, let us, have, am, will not, would, is, not]

  • Remove unnecessary words and POS

[ok, yes, no, hmm, yeah, uh, huh, to, Um, Oh, Alas, Oh, Eh, er, uh, uh huh, um, well]

[Article, Preposition, Conjunction, Determiner, Modal, Interjection, Numeral, Punctuation]

  • Tag the user’s input with its Part of Speech (POS).

  • Tag the VPF match of the user’s input with its Part of Speech (POS)

  • Filter both input and VPF match, based on a list of the global keywords returned by the VPF Web Service.

  • Compare what is left for POS and values. For example, for my question “Does the castle has any other gates?” only the (gates castle) keywords are returned.

  • If the comparison is successful allow the output (i.e., a script containing all the synchronized animations, speech, etc) of the VPF service to be executed by the system

  • If the comparison fails, pass the input to Stage 2 for deep syntactic processing.

    Stage 2: Deep Syntactic Processing

  • Fully Parse the user’s input and extract its predicates and Deep Dependencies (e.g., Subject, DirectObject, etc). For example, my phrase “I would like a brief description about all walls!” fails in the first stage of processing and parses like: 

like(Subject: I, DirectObject: description, SpaceComplement: about walls)

  • If it is a single predicate sentence, conduct 10 similarity tests, to check for similarities between the parsed input and the pre-parsed sentences in the database.

  • If a match is found, query the VPF with the match.

  • Return the output (i.e., a script containing all the synchronized animations, speech, etc) and execute it in the system.

  • If it is a double predicate sentence, conduct 9 similarity tests, to check for similarities between the parsed input and the pre-parsed sentences in the database

  • If a match is found query the VPF with the match. 

  • Return the output (i.e., a script containing all the synchronized animations, speech, etc) and execute it in the system.

  • If this stage fails, ask the user to rephrase or to move on to another question.

    I also experimented with the semantic parsing, but the Antelope parser is currently experimental. I am planning to add a third stage for full semantic processing in the current algorithm once Antelope’s parser is fully matured.

Friday, May 7, 2010

Videos – Prototype 1 (Part 2 Greek Only)

This is the second part of the vide for prototype 1. The system has just completed the presentation for location 1, and asks the user to move on to location 2.  Notice how the agent reacts at the beginning of the presentation (she can not see the user through the camera). The reaction is similar to the one you would expect from a real guide, if she saw that someone from the group is not listening to what she is saying.

Videos – Prototype 1 (Part 1 Greek Only)

This is the first part of the two videos for Prototype 1. The guide’s task is to provide navigation instructions on predetermined routes, as well as personalised information for specific locations in the castle of Monemvasia.

When the system loads the user is given a choice between three information scenarios: Architecture, History and Biographical. The user can also customize the appearance of the agent and other system settings, but this doesn’t show on the video. During a presentation the agent can utilize information from the Face-Detection module, and react if for example, the user is standing too far away from the camera. Finally, notice the use of FSM (Finite State Machine) in the construction of the dialogues. With the proper authoring tool such dialogues are extremely easy to make and can cover a whole range of dialogue phenomena.

Another idea I experimented for a while was the use of emotional responses as a method to guide how a presentation evolves about a location. For example if the user is too bored of the provided information, the agent can try to either provide alternative information or speed up the pace. However, its impossible to implement such approach in the existing Script-based systems. Possibly a KB is needed to dynamically create the contents of each presentation, but how the agent augments the story with non-verbal behaviour is an open question.

Thursday, May 6, 2010

Videos – Prototype 3 (Clothes)

This is the same prototype system (Prototype 3) but with a clothed avatar. The clothing is completely dynamic and very demanding in hardware resources. It won't even open on the UMPC. The run the clothe animations smoothly (along with the avatar ones) you need a quad-core system. In my opinion the avatar looks more realistic with the clothes on than without. The clothes are available free-of-charge at http://virtual-guide-systems.blogspot.com/2009/09/current-developments-haptek-clothing.html

Monday, May 3, 2010

MGUIDE Components

Below are a number of components I used in the development phase of MGUIDE. I am giving them away for free, all you have to do is to email me at virtual.guide.systems at googlemail.com.

1) AIML Control (.dll): This control can be added into any authoring environment (e.g., Flash, Director, etc) as long as you have the latest .NET framework installed. The control can handle Unicode characters and has full Javascript support.

2) WebCam Capture (.dll). A simple control that allows to add web cam support into any project.

3) Face Detection (.dll). A control that allows you to add face detection into any project. It uses the Intel’s OpenCV library for face detection and a .NET wrapper from Mr.Oshikiri. The control can detect right, left, far_away, normal and close. It currently uses OpenCV 1.0, and if you want to upgrade it to the latest OpenCV 2.0 you will need the latest .NET wrapper from here   

4) QRCode (.dll). The control allows you to add QR-CODE recognition into any project. It uses two commercial components from a company called PartiTek.

a) PtImageRW.dll

b) PtQRDecode.dll

You will need to purchase these components if you want to make the control to work.

5) Subtitles (.dll). A control that allows you to add speech subtitles to any project. The control uses a number of commercial components from Chant.

a) Chant.Shared.dll

b) Chant.SpeechKit.dll

c) DNSpeechKit.dll

d) NSpeechKitLib.dll

and an XML file as a subtitles feed. You will need to purchase Chant SpeechKit if you want to make this component to work.

Finally don't forget my free offering on Haptek Clothing at: http://virtual-guide-systems.blogspot.com/2009/09/current-developments-haptek-clothing.html

 

Tuesday, September 29, 2009

Current Developments - Haptek Clothing

After spending several months building these clothes I realised that they are too complex to be handled by the limited hardware of the UMPC device. Hence, I decided to give them for free. The clothes come with several textures, animations, morphs, etc. The clothes are designed for the character(called the guide) I use in my systems. I will possibly release the character once I am done with the evaluation stage of my project.

You can find them here:

Monday, September 28, 2009

Current Developments - MGUIDE Prototypes

Prototype 1:

 

1. A 3D Agent with more than 2000 gestures and several face expressions. The agent uses this body and facial language to augment the location presentation and navigation instructions provided.
2. A 3D agent that is aware of its environment and can dynamically evoke the attention of the user during a presentation for a location. For example if the user is poking around for a certain period of time, it can request the attention of the user to the presentation.
3. A 3D agent with fully dynamic 3D clothing with changeable textures during system configuration from the user.
4. A 3D agent capable of using additional multimedia information (on a 3D board) to enhance further the transmitted information (mainly in the presentation mode).
5. A Finite State Machine (FSM) dialogue manager capable of dynamically displaying questions based on the user’s selection and the current context. The questions cover a very broad range of the possible questions/clarifications that a user can ask after a presentation for a location.
6. 12 information scenarios based on what the castle has to offer to the potential visitor both culturally and historically. The total content (presentations and questions) is more than 10 hours long.
7. Designed but not implemented, customization of the agent voice.

Screenshot 1: The animated agent points to an image on the 3D board


Screenshot 2: The animated agent gestures as she speaks

Prototype 2:

Similar to the first system but with one additional feature - QR Code based navigation. A QR-Code is a bar-code capable of storing up to 4,296 characters in a simple geometrical shape. The system uses a QR-Code recognition algorithm to recognize the locations that the user is currently in (the user must photograph the QR-Code in order for the system to process it).
 

Prototype 3:

Similar to the other systems but it focuses only on the provision of navigation instructions. At the moment the system uses only photographs of landmarks but other more automated methods (e.g., GPS positing) have also been considered.
 

Prototype 4:

It features one information scenario only, along with the characteristics of the first prototype and:
2) Dynamic changing voice recognition grammars that allows the user to interact with the system using only h/er voice.
3) Natural language processing abilities using the Stanford/Link Parser. The system utilizes a novel algorithm that allows it to conduct predicate analysis and score keyword matching (if the first stage fails).The second stage of analysis is conducted by a secondary web system (Virtual People Factory).
4) A highly experimental search/comparison algorithm utilizing Semantic interpretation of the user’s input.

Screenshot 3: The system preferences of prototype 3 (Natural Language Processing version)