Should Professionals Be Using ChatGPT at Work?

Increasingly, many of us are using ChatGPT and other GenAI tools* for work to help with a diversity of knowledge tasks. We may share with our colleagues how much doing so has improved how we work, for example, saving us time and making us more efficient, while revealing the ways they are helpful, for example, generating ideas, writing or seeking information. However, in many workplaces, it is frowned upon to use them and in some outrightly banned. The main reason given is confidentiality; the need to protect sensitive, personal or proprietary data. Many banks, tech companies and healthcare providers are concerned about the risk of data exposure, where workers might divulge confidential information (e.g. new products or plans for new investments) when using ChatGPT. Instead, they are provided with in-house AI tools, which may not be as good or easy to use.

It is a well-known secret, however, that a number of professionals, who work in organisations where public GenAI tools are banned, may use them furtively on their own devices, while working from home, and without letting their colleagues know. Recognising that the genie is out of the bottle has led to some organisations rethinking their policy on barring the use of ChatGPT. For example, earlier this year, the UK’s Department of Work and Pensions (DWP) reversed its ban. It has recently begun allowing its civil servants to use them for official business or when using government-issued devices (with the exception of DeepSeek). They realised the benefits outweighed the risks. On their website, it now says how it can help employees respond more quickly to queries, while providing their customers with “a more personalised and seamless journey and access support how and when they choose”. Not only does this public acceptance remove the stigma and guilt of using ChatGPT at work for these employees but it might end up triggering a snowball effect, leading to other departments following suit. The question this raises is where should organisations, who have sensitive data and proprietary information, draw the line for what is acceptable practice and what is not?

Consider the healthcare profession where it is important to get this right. Similar to DWP, it is now widely accepted that using ChatGPT at work can be useful and beneficial for clinicians, especially as they have lots of admin tasks, such as note-taking, generating summaries, and treatment plans, which we know ChatGPT is very good at. Most likely, most people, including patients, would not object.

But what about other aspects of general practice where clinicians would like to use it, but where patients might find it unacceptable? To find out, the UK’s NHS is trialling various AI tools in a number of GP practices, where they will be used to summarise consultations with patients (online or in person) that will then be used generate clinical notes and referrals using customised templates. Findings from an initial pilot study run at GOSH, using the ambient AI voice technology called Tortus showed it saved a lot of clinician time. Patients were also very positive, noting how it enabled the cllnicians to give their full attention to them during the consultation.

Where using GenAI tools is clearly unacceptable is when making a decision about a patient’s treatment or surgery. It is considered a no-no for clinicians to ask it to suggest the best course of treatment for one of their cancer patients. What would a patient think, if a clinician said to them, “I just checked with ChatGPT and it recommended we use a new form of cryotherapy that freezes the cancerous tissue within your prostate to destroy the cells rather than opting for the more commonly used High-Intensity Focused Ultrasound (HIFU) that heats and destroys cancer cells within a targeted area of the prostate.” Most would balk at the idea that a machine was deciding their fate. Patients in this situation are highly anxious and want reassurance that the decision that is being made about their treatment is being made by human expert doctors. Even though the AI could be trained to make better and more informed decisions like these, patients will most likely persist with wanting human reassurance and human decision-making. 

What other aspects of clinical decision-making might patients be more willing to accept for the AI to do if it helps the clinician in their work? What about the research clinicians need to do to keep up to date and discover the latest procedures? Rather than looking up information in online medical journals (e.g. PubMed), themselves, why not ask ChatGPT? It could speed up the research process for them in the way many of us now regularly use ChatGPT to get started on a project. It might also suggest alternative types of surgery or care plans that the clinician might not have thought about or discovered by themselves. Would this use of ChatGPT as a research assistant be acceptable by the medical practice, if it was increasingly found that clinicians were already doing this but without letting on?

Besides the various ethical reasons that have been espoused for not using AI at work, human nature, itself, can play a role in determining what is acceptable and what is not. Our propensity to judge each other all the time about what we do, what we eat, what we like, our appearance and so on is also shaping our perceptions of whether it is OK to use GenAI. Some people will shake their head in disapproval when discovering one of their colleagues has been using ChatGPT at work for tasks that were previously ‘done by hand’, for example, using it to write a farewell note for someone’s leaving card, composing a welcoming speech for new employees, or summarising feedback following an appraisal. It seems disrespectful for those on the receiving end, to the extent, some of us might feel shame if we were found out to have done this.

Another growing complaint is the output from AI tools is bland, homogenous, and lacking personality. A new term that has caught the public’s imagination is sycophancy, which refers to how AI appears to always want to please the user, agreeing with them, giving them positive feedback while avoiding providing criticism. To be human is often far from being sycophantic – we like to be different, funny, critical and at times we can’t help being sarcastic – qualities that AI has yet to demonstrate in any human-like way.

So, where do we draw the line between what is acceptable and what is not when using GenAI for work? As new versions of AI tools materialise that are smarter with more safeguards in place, it will probably end up being a case of moving the goalposts. Furthermore, as people start using them for a wider range of work tasks, it may have the knock-on effect of changing their opinions and perceptions; they may become less concerned about professionals (e.g. teachers, doctors, lawyers, financial advisors, pilots, and politicians) using them in their work.

Current research is also discovering there is a shift in professional’s perception of using AI to help them with their work.  For example, radiologists have begun to value more AI assistance for certain tasks such as gathering relevant data to inform their decisions. Pilots, likewise, are also more willing to have AI at hand to assist them for certain tasks; for example, presenting seminal information clearly and highlighting constraints at nearby airports. They draw the line, however, for when the AI suggests a course of action they should take. That is their prerogative.

In sum, so long as the role of GenAI is to assist, i.e., to inform us, enable us to complete our tasks more effectively, suggest alternatives we might not have come up with ourselves or extend what we can do, then it will continue to be increasingly adopted in all manner of workplaces. It is only if it starts to be used to take over complex and sensitive human tasks that it will be concerning and troubling, Many people will very likely want humans to continue to do them. 

* I use ChatGPT here as an umbrella term for all GenAI tools. The image was generated using ChatGPT-5

Sleeping Patterns

One of the research projects I am currently working on is exploring how a wearable technology – the Oura ring – can help older people learn more about their sleep. We are conducting a study to see how older adults (over 65 years) take to wearing them. Will they find the data collected insightful or too much information? What will they pay attention to most (there are many data parameters that can be looked at)? Will they change their sleeping habits in an effort to improve their health?

Many of us go to bed in the hope we will nod off and wake up 8 hours later refreshed and ready to face a new day. However, as we get older we sleep less and more fitfully. I tend to sleep about 6-7 hours a night whereas when I was young it was more like 8-9 hours. I also wake up more through the night and sometimes have to make a conscious effort not to think about things that are worrying. A motivation for our study is to explore how older people manage their sleep and whether this kind of intervention is seen as being helpful.

To gain initial insight into how wearing the ring may increase our awareness and even improve our sleeping habits, myself and my fellow researchers on the team have been wearing one each now for a month. We have kept a diary. My entries were daily for the first two weeks, full of observations and then slacked off. Here are a few from the beginning.

Day 1 Monday 20th Jan 2025
Arrived in the afternoon in a white box that was aesthetic to look at and very easy to open. It felt like an Apple product. It was easy to try on and to download the app and get it set up via Bluetooth. The instructions for wearing it were nice and simple. Much thought had gone into the opening and onboarding experience. I looked at the visualisations on the app for resting heart (it was quite high as I was excited to try it on). I then saw it had gone down to a reasonable score in the evening.

Day 2 Tuesday 21st Jan 2025
I felt that I did not sleep well last night, waking up a few times but the visualisation from the Oura app said I had had a good night’s sleep. It showed periods of shallow, deep and REM sleep. It also said I fell asleep in 7 minutes which seems about right. I fall asleep quite quickly most days. I slept for 6.33 hours. And have a sleep efficiency of 82%. My sleep score (whatever that means is 83 Good). My readiness score was also rated as good, too, whatever that means. Numbers, eh.

 

Days 5 and 6 Friday 25th/Saturday 26th January
I went to bed at midnight and did not sleep well. The app data reflected this saying I had had only 5 hours sleep and not a good night. So the next night I went to bed early and got 7 hours sleep. I got a sleep score of 89 which equates to being optimal. And a sleep efficiency of 89%. Back on track. Just goes to show how resilient we are bouncing back if we rest up after a late night.

Then two days later:

Monday 27th January
I had a great night’s sleep last night. I got the highest score of 92 with an optimal rating and a crown icon. Total sleep was 7 hours and 23 minutes, good everything else. Text all in blue with no reds. The ‘in focus’ message that accompanied my high scores was surprisingly effective at making me smile. I must be a sucker to gamification. “Your 7 hours 23 minute of sleep last night will help you think sharp and stay focussed. Have a great day!”

It seems this yo-yo sleep pattern (poor night, poor readings, good night, recover, good readings) is part of my lifestyle. The data from the ring has certainly woken me up to this but will I change my lifestyle? Go to bed at 10.00p.m each night after a cup of camomile? I doubt it, not just yet. Maybe when I get older. If the body is so resilient why try to optimise it every night? Or is it even possible?

The new Bridget Jones film “Mad About a Boy” has just come out. The first film, “Bridget Jones’s Diary” (2001) was about a 32 year old single woman who kept a diary to improve herself. I wonder if it helped her to do so and what she does now she is 20+ years older.

Super Shoes

Kelvin Kiptum Nike shoes Photo by Michael Reaves/Getty Images)

This autumn, Ethiopia’s Tigst Assefa broke the woman’s world marathon record. She took just 2 hours 11 minutes and 53 seconds – which is 2 minutes and 11 seconds less than the record set previously 4 years ago. That is a whopping amount of time she was able to shave off. Not surprisingly, it raised eyebrows. How was it possible to run so fast? Some commentators put it down to the trainers she was wearing that gave her the advantage – the Adizero Adios Pro Evo 1. Not long after, the Kenyan long-distance runner Kelvin Kiptum broke the man’s world record time, wearing the latest Nike Alphafly 3 trainers (see left). So, what is it about these new kinds of super shoe that literally make an athlete run like the wind?

A big step change is the way they are made up and the materials used for this. This has enabled a new thick but lightweight structure to be built in the sole of the shoe. The way the layers of material are engineered seems to give the runners that extra spring. They also have added a stiff rod in the midsole that is made of carbon, which helps the shoe keep its shape. The Alphaflys also have a curved geometry in the sole that has been designed to propel runners forward. Taken together, they are truly a step up from previous running shoes. A marathon runner who was interviewed in an article in the Guardian on the super shoes said “on average I reckon that they are worth four minutes for a top male in a marathon.” And the proof is in the tumbling records this year.

But is it fair?

Technology has been developed for years to improve all manner of artefacts, clothing and equipment that are used in sport – including tennis racquets, cricket bats, racing cars and racing bikes. It is par of the course in sport innovation. But some ruling bodies see it as unfair and needs to be stopped in its tracks. For example, back in 2009 a new kind of high-tech super swimsuit developed by Speedo was banned by the swimming’s governing body on the grounds that it gave certain swimmers an advantage, and in so doing, was ruining the sport. The full body swimsuit was made from polyurethane which can trap air in the suit and increase buoyancy to the swimmer making it faster to swim.

The LZR Racer Suit unveiling at a press conference in New York City in February 2008. From wikipedia

Another concern is its impact on the past. Many world records from years gone by are being broken, especially those that were made by great sportsmen and women, who have since become legends. They did not have the same kind of high-tech super shoes, etc., then, so it is considered unfair to take away their crowning glories by those who have the super shoes, swimsuits, etc..

But records are there to be broken. And speed and innovation go hand in hand.

Another line of argument is that giving sportspeople these new kind of superpowers is equivalent to doping which entail taking certain kinds of drugs to improve fitness. But the big difference between doping and super tech is that the former is an internal enhancer while the later is an external aid. While both can improve speed, doping goes one step further by invisibly enhancing an athlete’s stamina which is especially important in endurance and long distance sports. The various substances, like EPO, used in doping increase the taker’s red blood cell count, enabling more oxygen to be transported around the body and to the muscles, thereby increasing stamina. No-one questions whether this should be banned because not only can they give certain athletes an unfair advantage, they can be dangerous, causing health issues. It is also difficult to see how much someone has taken and for how long so it is a very unfair playing ground.

Super shoes, on the other hand can be checked to see if they fit certain regulations. The same could potentially be true for full body swimsuits (they still remain banned from the ruling of 2009).

Ten years ago Google developed a concept shoe that could talk to the wearer to motivate them to get up and go. It was long before chatGPT had arrived. In the near future it might be the case that the super shoe could be embedded with a GenAI app so that it could talk to the runner like in the video – helping them keep going in the way trainers and spectators currently do from the sidelines. Now that would be truly super.

 

Oh Bard!

Sometimes bad timing and misfortune can end up having a massive negative impact on an organisation, as happened recently to Google who were in the process of launching their new AI tool, Bard. Hours before going live Reuters pointed out how it was not up to scratch as it saw an error in the promotional ad. It went viral and the effect was to wipe billions of dollars off Google’s shares. How did it happen?

A tiny factual error in one of Bard’s maiden answers to what was a seemingly banal question was the trigger for this catastrophic nose-dive. The question in question that Bard was asked was what new discoveries had been made from one of NASA’s mighty big space telescopes (the JWST) that could be told to a nine-year old. Bard replied that it took a picture of an exoplanet – which is a planet outside of the earth’s solar system. However, the human who tweeted pointed out that in fact it was another telescope that did this – one in Europe called a very large telescope (the VLT). It was meant to be an answer suitable for a young inquisitive child and most kids of that age would have not minded or would have blurted out it had made the mistake.

However, a bit of investigative work by a reporter at the Financial Times noted how Bard was technically correct since it was the JWST’s very first sighting of an exoplanet, but in the wider context of world knowledge, it was another telescope that had spotted it earlier. So a pedantic matter. Just goes to show how fickle the world is when it comes to its trust and faith in tech. Or maybe it was fuel thrown at the new AI race between Google and Microsoft.

Meanwhile OpenAI’s chatGPT (with Microsoft investment) continues to soar in terms of its credibility, popularity and capability. I have used it several times now and am amused and astounded by what it can accomplish in real time. Sure, it can get things wrong (for example, it did not know the Queen had died or when the King’s Coronation is because it is only trained on data before 2021) and its prose can be a bit bland and clunky but it has transformed how many pedestrian writing tasks can be achieved. Just like the spreadsheet changed how we do financial forecasting and the calculator offloaded the need for humans to do mental arithmetic in their heads anymore so, too, will this new generation of LLMs transform how we write.

In fact, millions of people, like me, have tried using ChatGPT in the last couple of months and are mightily impressed by how it can get them started writing an essay or report – overcoming that blank page syndrome. When I asked it to write some feedback that I could give for a graduate student report I was impressed by its fluid style and use of praise – almost as good and personalized as I could!

At the same time there are those who are worried that it will dumb us down or turn us into cheats, for example, students will increasingly use it to write their essays, reports and other assignments on their behalf. But why not? They can then be asked to read and spend time reflecting to how good ChatGPT’s answer is and how they can improve upon it. Instead of simply regurgitating what they find on Wikipedia or other online resources they could be asked to develop and hone their critical and analytical skills. And learn what makes for a good or poor argument, developing some metacognition skills in the process. Meanwhile, professors and teachers could use the next generation of ‘turnitin’ AI plagiarism tools that are starting to appear to detect how much they have changed the chatGPT answers. We can also begin to rethink our assignments and ways of providing feedback to students. In so doing, we can all learn to write better – be it generating and creating or assessing and providing feedback. Framing the new generation of AI in this way will enable all sorts of new possibilities for students (and teachers) to learn and teach with. As was said in the Google launch blurb Bard “can be an outlet for creativity, and a launchpad for curiosity.”

Funny how scientists love coming up with acronyms so much. Anyone want to guess what NASA, JWST, VLT, LLM and GPT stand for? Perhaps we could just ask Bard.

Numbing numbers

A part of my work involves advising various companies and government bodies about their future technology research strategies. This week, I had the pleasure of visiting the EPSRC headquarters in Swindon for a two-day meeting. Besides being super excited about meeting up with other members of the scientific advisory team for the first time in 3D (many of whom I had only met as digital postage stamps during the past 3 years), we were also invited to tour various facilities at the Harwell Science and Innovation Campus in nearby Didcot. I was assigned to a small group that went to see the Diamond Light Source. Sounds like a big laser show but it is in fact a massive science centre. On entering the building we were shown a model of the building to see what is inside. From an aerial view, it reminded me of Apple’s latest donut campus – in terms of its shape, scale and size.

This huge set up enables a laser beam to whiz around a circuit at some unbelievable speed that has the effect of giving off light 10 billion times brighter than the sun. The mind boggles at what that actually means in practice. Of course you never get to actually see the beam racing around the donut but I was mightily impressive to know it was hurtling around while we mere mortals were walking tentatively around the place. What you do see are various numbers, graphs and images in the nerve centre that lets you know what it is up to.

 

 

 

 

 

 

The shining beam is siphoned off into lots of little labs where teams of scientists use it to study all manner of topics, including fossils, jet engines and vaccines. In a nutshell, it is a ginormous microscope that allows lots of measurements to be taken during millions of experimental trials. The scientists and engineers we were introduced to, who work on site, told us how having such a powerful magnifying beam enables them to do incredible science. Making the invisible visible in all sorts of unknown ways. However, they all also mentioned that the data that is collected from the various beam runs is so massive, that outside researchers who come to run their experiments at the facility, nowadays can’t cart their data home with them because there is simply too much of it. Instead, the data has to be stored and analysed on site before a digested form is sent back to them. A few years ago they could save it all onto a portable hard drive…

So what do scientists do with this ever growing mountain of data? How do they count it, describe it and then importantly, work out how to turn it into new knowledge? While data scientists may revel in this sea of data, other kinds of scientists often find it somewhat overwhelming, to put it mildly. To help them in their quest, every now and again, the ‘numbers’ community come up with new names to  make it more manageable to describe it. Just this month, the 27th General Conference on Weights and Measures introduced some new names (or prefixes, cf. milli- to be precise). These included talking about data in terms of being a ronna which has 27 zeroes after the first digit or a quetta which has 30 zeroes after the first digit.

So how would you use these funny number names to talk about science? Consider the mass of the sun. Previously it was talked about in terms of 2,000,000,000 yottagrams. Now its mass can be said to be 2,000 quettagrams. To me it all sounds a bit double-dutch but to data scientists it provides a new and hopefully more meaningful way of expressing huge quantities of digital information.

Coming up with new data words must be one of the fun parts of the job. I wondered how did they choose them this time?  What were they thinking of? Ronna is also a girl’s name meaning “rough island and true image”, while a Quetta is a fort and is the tenth most populous city in Pakistan! And what will replace the ronnas and quettas to express the ever growing volumes of data that will have accumulated in the next 5-10 years? What new words will be dreamt up?  Perhaps Tom Lehrer should get in on the act. I am reminded of his classic song, The Elements, from 1956. Makes me laugh every time I hear it.

Thanks to John Collomosse for taking the photos inside the Diamond.