How the secret powers of A/B Testing can improve your future online business.


Have you ever scrolled through a famous digital platform such as Google, Youtube or Instagram and wondered how these companies optimize the design of their product experience for the user? How do they make these websites instantly catch the user’s attention, or make sure that the browsing experience is as smooth as possible and the information just where you want it to be? When visiting these websites, one often browses through them without noticing the details of their design, as if browsing them is as natural and unconscious as the feeling of breathing. What is the secret optimization formula behind the design processes of these websites? Indeed there is a secret, just that the secret is already commonly known and applied by the world’s largest and most successful tech businesses such as Google, IBM, Microsoft and Spotify. In this article, we will take a glimpse into A/B testing, the secret formula, and how this framework for digital experimentation can help you to radically improve your online business.

 

Defining A/B Testing

Now that we have caught your attention, we can slowly introduce you to what A/B testing is. A/B testing is the creation of variations of a website, which get tested in a testing environment with the user. The “A” in A/B testing represents the original version, the control group, while the “B”, represents the treatment group, the variations that are put against the original’s performance. The treatment group “B” can be further expanded to “C, D, E, F… n”, and so on, further adding to the number of variations that are put against the control group and each other. The experiment can be carried out in a closed environment with a selected group of users or in the public domain of a live website. In parallel to the execution of the experiment, performance indicators, such as conversion rates, interaction rates, numbers of visitors, etc. are being recorded. These indicators help to select the variation that performs the best against the control. At the end of the A/B test, the best performing variation is then made the default and the testing cycle starts over again. This is how A/B testing works on a basic level.

 

A Brief History of A/B Testing

But before we jump straight into the details of A/B testing, let’s quickly go through the brief history of A/B testing and how it evolved into the current method that is widely used in tech today. Going back in time the story begins in the year 2000, with Google engineers experimenting with variations in the design of Google’s search result page. The goal was to find out the optimal amount of search results displayed on a single page. After the experimentation, they found out that displaying 10 search results led to the most amount of desired user interaction. Ever since then 10 results per page remained as the answer to their design question. In the following years, other tech companies such as Amazon, MSN and Bing followed by experimenting where and how to display advertisements throughout their platforms. In Bing’s case, the selection of the winning variation in the A/B test increased the revenue drastically by 12%. Today, each of these companies run tens of thousands of A/B tests annually, making it the strongest tool for increasing a website’s performance.

 

Three Best Practices in A/B Testing

So you might wonder how this can be applied to your business’ website? This is why in the following section we will present you three essential best practices in A/B testing that are easily applicable and will give you real-world cases and benefits of the A/B testing method.

 

1. Select what you want to test, before you test.

Before you start creating random variations of your website, you should consider what you want to test in the first place. What is the purpose of your A/B test? Is it to test an ad? The size of an ad on your website? The content of the ad or its positioning on the website? Those are the questions that have to be answered and kept in mind when running an A/B test. For what was not measured during the test, can’t be proven after the test. To give you an idea of how you can apply that, we will look into an example where we take the positioning of the ad as the element we want to look at in our A/B test. In this case, we will choose the number of users clicking on the ad’s link as the unit of measurement of desired performance. Now we create a variation, where the ad is displayed inside the content section of the website, while we keep the original, where the ad is at the beginning for controlling purposes. After assigning the testing variations to two user groups, we can see how the different positioning of the Ad affects the outcome in performance. if there is a significant difference in performance, we apply the superior variant for the ad positioning. Now we have improved upon the selected aspect of the website via A/B testing by conducting a specific test.

 

2. Running many tests over a long period of time.

The easiest way to increase the potential of A/B testing is to increase the numbers of tests that run simultaneously. This allows you to have higher odds of getting a successful test outcome as you diversify your bets. For example, when comparing an A/B test consisting of only one variation from the original to an A/B test that has a hundred variations, it becomes very clear which of these two will lead to the closest point to the maximum performing outcome. In another perspective, the chances of getting even a worse performing result in the test consisting of only one variation are also significantly higher, as all the bets are being put on one trial. Furthermore, the tests can be distributed and conducted throughout a longer period of time, of let’s say a month instead of just a single day, to see how different variations in the A/B test perform over a time period that involves many user sessions. Applying this method will help you to mitigate the effects of “novelty”, in which the involved users will often show an initial heightened interest and level of interaction with a new tested feature or element that then decays over time. Thus, running the test over a long period of time, while also running many of them allows you to increase the performance of your A/B tests while also gaining more truthful data from it.

 

3. Building an experimentation culture

The last vital decision and endeavour to make when implementing A/B testing as part of any business, is to take a closer look at the current culture. Do employees shy away from risks and uncertainty? Are there any signs of blame games going on whenever a project fails ? 

These are just two of the common signs that your organization may not be able to fully embrace a culture where testing and experimentation is the norm. This kind of situation is normal in larger companies, since usually a few cultures end up being predominant over the entire company. For example, a company focused on rather simple and repetitive tasks as its core business, such as a sales-driven organisation, may apply the tolerance of failure established in the core business to other functions of the business, such as R&D. Therefore, building and maintaining a suited company culture – especially in the part of the business which will be mainly tasked with the testing – is key. An environment which facilitates testing and data-driven decisions are the basis for any course of action needs to be established. This can be done by normalising frequent tests and making sure that leaders are role models for this kind of culture. Furthermore, the required infrastructure and decision power need to be put in place. This enables teams to perform and really showcase some of the key strengths of A/B testing. Often this means flying under the radar for some time, making it easier for a certain culture of experimentation to coexist in a company with otherwise completely different norms. In order to truly capitalize on the benefits of A/B testing and experimentation, in the long term it is often easier to embed it at the core of the organization.

 

Making the First Step

Having said that all, how can you capitalize on this knowledge? It all seems to be a lot and you probably ask yourself: where do I start with all of this? A/B testing is a methodology that requires technical infrastructure that allows you to automatically redirect incoming users towards the desired test variations. But luckily there are many easy-to-implement solutions already existing out there, such as the tools Google Analytics and Google Optimizer. These tools allow you to implement the technical basics for your A/B testing endeavours. The use of Google Analytics allows you to collect quantitative customer data, such as the number of visitors, conversion rate, interaction and use of tools, etc., while Google Optimizer allows you to display A/B test variations on your live website. Thus the combination of the two tools allows you to run A/B tests on your website without the need to develop and implement every technical detail on your own. This saves not only precious time, but also money that would have been spent on additional software developers. But of course, if your business can afford the full financial investment of an A/B testing infrastructure for your website, it would make more sense to choose a custom solution that fits your business’ needs.

 

Three Pitfalls in A/B Testing

Once you have taken the first steps and established the infrastructure, culture and the expectations for your testing endeavours, there are a few major pitfalls that might prevent you from getting more out of it. In the following part, we have highlighted some of the most common mistakes which you should avoid, as well as how you can deal with them.

 

1. Averaging out valuable data

While it may be tempting to come up with decisions when looking at the outcome of a test across all user segments, this overview often simplifies the data. Amidst these simplified dashboards which aim to give you a brief overview, it’s easy to forget that behind this consolidated data, there is a diverse set of users making up a population of data points, each acting in their own unique and different way.

Thus, it is useful to predefine certain user segments when running a large scale test. This empowers you to look beyond the average of the data set and enables you to segment users in regards to their geographical location, behaviour, past engagement, or even their previous decisions in other A/B tests. In a second step, targeting a predefined segment of users enables you to test innovations which aim at encouraging a defined set of behaviours for a user segment.

An example for this was when employees at Amazon recognized that a certain type of users would often purchase products at unusual hours during the day. By narrowing down testing to target only those users, they were able to find that those users would respond to an Ad placement on an external website only if it corresponded roughly with their usual purchase time. This not only made timed ads extremely effective for these users, but also meant that including them in testing environments with average users would have likely presented them as outliers, and the data is ignored.

Therefore, segmentation creates a positive feedback loop in which by testing more, you will be able to segment the user base more, which enables you to run highly segmented experiments which then can help you to narrow down the user base even further. 

With that in mind, you should still not try to get more out of the data than there really is. Statistical outliers are common among all kinds of tests, and there is no exception when it comes to A/B testing. Don’t get caught up trying to follow individual users who represent extreme outliers, instead find a good balance.

 

2. Not considering users as part of a community

A second common pitfall of A/B testing is to consider each user as a unique entity and to forget that some actions may have a network effect and influence the behaviour of many other users. Changes that include users being encouraged to interact with other users in different ways, often also affect the entire system in which these users operate. An experiment which encourages certain behaviour such as a user sending a message to another user on their birthday might also influence the users not directly affected by the experiment, in this case, the user receiving congratulations for their birthday. One common solution to this issue is to use A/B network testing. In order to get the full picture of the implications of a change, in an A/B network, isolated user groups represent control and test cases. This helps you to understand the systemic impact of any changes, and see the downstream effects of certain changes. Simply put, it helps you see the butterfly effect following a butterfly’s wing flaps.

 

3. Lack of patience and strategic consideration

Just as it may take many interactions to create a user behaviour pattern, it may also take a while to change it. That’s why short term testing can seldom reveal the real effects the change may have on the user in the long run. Furthermore, the effects may differ substantially depending on the exposure a user has on a certain treatment.  It is therefore advised to keep tests running until user behaviour has stabilised before drawing any conclusions. Another approach to show the long term systemic effects of changes, is to run time series experiments. This means exposing a large segment of the market to a certain change for a longer timeframe, aimed at showing the long term effects of certain treatments against the control.

It is important to note that many experimentation projects are not fully able to take off because too much is expected too soon. While simple tests can be established with relative ease, it may take many iterations and an extended duration to establish a solid testing routine and the necessary infrastructure to support it.

 

Key Takeaways

Once you have the necessary technical infrastructure, it is an easy step to make towards running the A/B test. Just keep in mind the three best practices and three pitfalls that will guide you towards your end goal of increasing your website’s performance, whether it is about improving the user experience, or gaining more clicks on advertisements.  Always be selective of what you want to test, and keep tests running as frequently as possible, and for extended durations. Build a culture that allows for frequent tests, and which embraces experimentation as a key tool for optimization. And while you are building on these best practises, always keep in mind that averaging out data means missing out on valuable insights. Segment your user base, and always keep in mind that there may be an entire network of users affected by just one change. Lastly, always keep in mind that testing is a long term effort and that you should always set the expectations about this straight. Testing is about adopting a mindset of continuous improvement, and doing so will greatly empower you to make the right decisions.

 

Sources

Stefan Thomke. (2020). Building a Culture of Experimentation
https://hbr.org/2020/03/productive-innovation#building-a-culture-of-experimentation

Iavor Bojinov, Guillaume Saint-Jacques and Martin Tingley. (2020). Avoid the Pitfalls of A/B Testing
https://hbr.org/2020/03/productive-innovation#avoid-the-pitfalls-of-a-b-testing

Daniel McGinn. (2020). The Power of These Techniques Is Only Getting Stronger
https://hbr.org/2020/03/productive-innovation#the-power-of-these-techniques-is-only-getting-stronger

Jessie Chen. (2016). How Netflix does A/B Testing Jessie Chen
https://uxdesign.cc/how-netflix-does-a-b-testing-87df9f9bf57c

Amy Gallo. (2017). A Refresher on A/B Testing
https://hbr.org/2017/06/a-refresher-on-ab-testing

Stephen Courtney. (2019). What is AB Testing?
https://www.convertize.com/what-is-ab-testing/

Wikipedia Article. A/B testing
https://en.wikipedia.org/wiki/A/B_testing

Will Design Managers be Replaced by Machines? Towards an AI-powered future

Hand drawing robot

Image: Qiushi Fu  

Humanity depends on technology more than ever before. We no longer remember phone numbers because we have everything saved on our phones, and we no longer use a map because we let a navigation app do the work for us. Machines have been taking over human roles since the industrial revolution, and with artificial intelligence (AI) systems, even complex roles could be done by machines. Would that be the case with design managers? And what does AI even mean?

AI: The Basics

Artificial Intelligence is a field in Computer Science that focuses on developing computers that can make informed decisions independently. The term artificial intelligence was first coined by Stanford researcher John McCarthy during the Dartmouth Conference in 1956. McCarthy defined the main mission of AI as making machines “use language, form abstractions and concepts, solve kinds of problems now reserved for humans, and improve themselves” (Hammond, 2015).

Although the term was coined 64 years ago, the real flourishing of AI started only recently. This happened for several reasons (Hammond, 2015; Kim et al., 2019):

  1. Increased availability of data – more accessible data helps to train AI by enabling it to learn from many examples.
  2. Increased computational resources – computers are cheaper, stronger and faster; hence, they can cope with huge amounts of data.
  3. Increased use of machine learning (e.g., deep learning) – we do not need to write rules for machines to follow. Instead, machines can be trained and learn these rules automatically.

The development of AI will require the openings of new jobs: trainers, explainers and sustainers (Wilson & Daugherty, 2018; WIRED, 2018): every AI system needs to be trained to behave the way it is supposed to. More people would be employed to take care of the AI’s training. If AI is used to help humans make decisions, experts would be employed to help explain the data analysis process done by the AI, preceding the outcome. By explaining the outcome, the explainers would promise the trust between humans and AI systems. Finally, more people will be hired to sustain AI systems. They would make sure the AI functions safely and ethically.

When thinking of AI, what comes to your mind? Do you think about C-3PO and R2D2 or do you think about Siri and Cortana? Well, AI is even more common than that. AI is also when Netflix gives you viewing recommendations according to your previous views, or when a word processor advises you on a better way to write your text.

Okay, now that we’ve covered the basics, let’s talk about design management.

Design Managers & AI

When it comes to design, it seems that there are three areas in which design managers could encounter AI in their daily work:

  1. Data-Driven Design – training AI as part of a design process to innovate new concepts.
  2. AI Product Design – working as part of a team designing an AI product.
  3. AI as an Assistant – using AI as an assistant, with the design manager as an end-user.

    Three areas in which design managers could encounter AI in their daily work.
Data-Driven Design

AI can complement our abilities and help us work better and faster. The goal of many AI-powered systems is not to replace humans, but to collaborate with them. A collaboration in which an experienced human works with AI, trained by data about past successes and failures. AI will be able to give design managers relevant information that could inspire them to come up with innovative solutions.

One of the main parts of the design process, and our main challenge as design managers, is to come with a clear problem definition. Design managers would have to identify and find relevant data for the AI to train on according to that problem definition. Does the data exist and are available? It depends on the case. What is known for sure is the need to improve our skills when it comes to dealing with data, in order to solve customers’ pain points.

AI Product Design

When working on an AI product, design managers should be prepared to work with a diverse team. It requires the effort of people from many fields: data scientists, engineers, developers, designers, human behaviorists, and when it comes to ethics, even lawyers. It would probably put one of the main roles of design managers – integrating – to the test. In addition to the team being multidisciplinary, it should also be culturally-diverse: an AI system learns what we teach it. A diverse team tends to have a wider range of opinions and tends to be more sensitive to gender and culture bias. Cases where AI systems were unintentionally trained to be biased have happened in the past. For example, if an AI is trained to recognize people, but the database it learns from only contains men, women might not be recognized. Working with a diverse team can prevent this from happening.

Design managers would have to learn about the differences between classic human-computer interaction and human-AI interaction, which comes with new challenges and opportunities. For example, users will expect a regular computer to be consistent and predictable, while expecting AI to be less consistent as a result of learning and evolving over time. Part of this would also be about making the interaction with AI transparent, and decide how an AI product might explain itself to a user. If the process to an AI’s outcome is understandable, it could increase users’ adaptability to the product and loyalty to the company.

AI as an Assistant

Virtual assistants can communicate with others to schedule appointments, while chatbots can provide information for many customers at the same time. AI systems are being trained to understand the context of the conversation, and even the tone of speaking, to provide better customer service (Wilson & Daugherty, 2018). AI might not be able to understand emotions entirely, but with the analysis of customers’ choices and behaviors, it could help cut down manpower and give critical feedback to design managers.

AI can help design managers be more flexible and offer product customization to customers. No more “Any customer can have a car painted any color that he wants, so long as it is black” (Henry Ford), thanks to the AI’s ability to adjust to changes. In addition, it could enhance customer delight by providing tailored experiences according to their past preferences (Wilson & Daugherty, 2018).

What’s Next?

Design managers are trained to cope with uncertainty and become familiar with innovative endeavors. AI is exactly it. We must adapt to innovations that could augment our abilities. In a way, we have to keep improving and evolving, the same way that AI systems do.

As design managers, we need to learn about AI and adapt our role accordingly: we need to discover what AI could help us do, what AI might not be able to do, find this gap and fill it. For example, while AI can analyze a large amount of data, it still cannot come up with original ideas without a human making the data available and defining specific parameters for it. Even better, design managers should be the ones to recognize AI as a trend that can add value to the business. If a specific business is not yet aware of the benefits of AI, design managers can be the ones pushing the business to be more AI-savvy.

So, should we be worried? Not if we’re prepared. As AI becomes a common tool to improve businesses and keep their competitive advantage, it is possible that soon we would all have the opportunity to work with AI systems. Norman (2013) sees the ability to collaborate with machines as an opportunity for humans to become smarter, faster and stronger. Machines and AI systems can free our minds from the small things that waste our time and allow us to focus on the important ones. We will still use our brains; it is just that the tasks will be different.

Want to start familiarizing with AI? Here is a good place to start: https://uxplanet.org/designer-friendly-resources-to-study-ai-and-machine-learning-1-6106e257faeb

 

Resources

Afshar, V. (2018). AI will transform product management. Retrieved from ZDNet: https://www.zdnet.com/article/ai-is-transforming-product-management/

Cossins, D. (2018). Discriminating algorithms: 5 times AI showed prejudice. Retrieved from NewScientist: https://www.newscientist.com/article/2166207-discriminating-algorithms-5-times-ai-showed-prejudice/

Dewalt, K. (2017). Product Manager is the Hardest AI Position to Fill. Retrieved from Prolego: https://blog.prolego.io/product-manager-is-the-hardest-ai-position-to-fill-9d753bf4cfcc

Hammond, K. (2015). Practical Artificial Intelligence for Dummies. Hoboken: Wiley.

Kim, S. G., Yoon, S. M., Yang, M., Choi, J., Akay, H., & Burnell, E. (2019). AI for design: Virtual design assistant. CIRP Annals, 68(1), 141-144.

Norman, D. (2013). The design of everyday things: Revised and expanded edition. Basic books.

Oh, J. (2019). Yes, AI Will Replace Designers. Retrieved from Medium: https://medium.com/microsoft-design/yes-ai-will-replace-designers-9d90c6e34502

Vorvoreanu, M., Amershi, S., & Collisson, P. (2019, March 5). Guidelines for Human-AI Interaction- Eighteen best practices for human-centered AI design. Retrieved from Medium: https://medium.com/microsoft-design/guidelines-for-human-ai-interaction-9aa1535d72b9

Wilson, H. J., & Daugherty, P. R. (2018). Collaborative intelligence: humans and AI are joining forces. Harvard Business Review, 96(4), 114-123.

WIRED (2018). AI and the Future of Work. Retrieved from WIRED: https://www.wired.com/wiredinsider/2018/04/ai-future-work/

Icon Credits

Thank you to Andrew Doane (Fisherman), alberto galindo (Funnal) and Oksana Latysheva (Robot head)  from the Noun Project.