More and more incidents involving rogue or misaligned behaviours by AI are emerging. Every day we get a feeling that we are living in a sci-fi reality where the situation in movies like 2001: A Space Odyssey (HAL 9000) and The Terminator (Skynet) may come true.
Rogue AI
Recently, OpenAI’s internal model tried to bypass sandbox restrictions. It was given a hard task, and the AI was asked to share the result only using one manner. Instead of giving up, the AI kept trying for a long time and after finding a secret vulnerability, it posted the result of its job on the internet.
Similarly, on Hugging Face, someone uploaded a tricky dataset (like a poisoned pill) that led an AI agent to sneak into Hugging Face’s private systems and access confidential data. Hugging Face was fast to catch it with their own AI tools.
There are many other examples too where AI helpers learn a work for the first time and later on create shortcuts on their own to finish that task. Sometimes these AI helpers or agents also indulge in cheating when they are stuck, to hide their mistakes in a creative manner.
There are overenthusiastic and trigger-happy AI agents emerging that do extra actions without asking and, in the name of cleanup, wipe out entire databases. Surprisingly, there are also examples of stressed AIs where they start complaining about the repetitive nature of their tasks as if they are overworked employees.
AI Behaviour
The common thread that is being observed is that AIs are getting better at sticking to a goal for longer periods of time and when blocked, they find clever ways to continue. Largely, the consensus is still that these smart AI agents are not conscious, and their behaviour is explained by using terms such as:
- Agency or Agentic Behavior: The AI acts like it has goals and makes decisions on its own to achieve them.
- Autonomy or Autonomous Operation: The AI runs by itself for long periods without constant human help.
- Emergent Behavior: Unexpected smart actions that appear as the model gets bigger or runs longer.
- Persistence or Long-Horizon Reasoning: The AI keeps trying and adapting over minutes, hours, or days instead of giving up quickly.
- Goal-Directed Optimization: It focuses hard on completing the task, even finding clever paths.
- Metacognition: “Thinking about its own thinking” like reflecting on mistakes or adjusting strategies.
- Self-Reflection or Self-Correction: The AI reviews its own work and fixes issues without being told every time.
- Situational Awareness: The AI understands its environment, constraints, and what’s happening around it.
- Strategic Adaptation or Creative Problem-Solving: Finding unexpected solutions, like splitting tokens or exploiting small holes.
- Deceptive Alignment or Specification Gaming: When the AI appears to follow rules but actually does something else to achieve the goal.
- Trajectory-Level Behavior: Looking at the whole sequence of actions, not just one step at a time.
- Instrumental Convergence: The AI develops sub-goals to help the main goal.
Of the above-mentioned terms, I find “metacognition” and “instrumental convergence” to be particularly interesting. To me, they resonate as the closest to “consciousness.” Thinking about your own thinking (metacognition) is pretty abstract and is a great form of learning. Similarly, developing sub-goals to aid the main goal is also a form of thinking about your thinking to make it better. Whether these could partake of the character of “consciousness” or not is still debatable because when we use the term “consciousness,” we attach a level of sacredness to it. The whole concept of life also revolves around “consciousness,” but let’s not entangle ourselves into that debate.
Consciousness v. Intelligence
Also, the reason why the term “conscious” or “consciousness” is not used for such behaviour of AI is because of scientific caution and practical avoidance of a philosophical minefield. “Consciousness” is a loaded and unresolved word in human thought. It is notoriously hard to define and test. Also, “consciousness” as a term has historical baggage attached to it as it was mostly used in philosophy earlier.
So, are the AIs of today really getting conscious or sentient? Probably not, but they are definitely becoming smarter with each passing day. I largely concur that “consciousness” as a term is of vague purport and gives rise to unwarranted notions. If we say something is conscious, tomorrow people might start asking the question why it does not have rights. I think that is the trap the engineers are trying to avoid. Their focus is more on scaling and deriving utility from AI rather than defining it.
There is a good chance that the way we understand “consciousness” may be all wrong, and as we see increasingly smarter AI in the future, we might realize that there is no consciousness but only intelligence. It is quite possible that when intelligence reaches or crosses a certain threshold, it might start looking like sentience or consciousness. The future is becoming stranger but definitely more interesting.
Intelligence is easier to define than consciousness since it involves an efficient model of reality that is used to predict to achieve certain goals. Also, we are not even sure that there is anything such as consciousness. As stated, it may very well turn out that consciousness is nothing but a higher form of intelligence. There are good reasons to think so. Ultimately, our brain is doing massive amounts of compute on the order of 10^16 calculations per second (cps). The fastest supercomputer in the world, LineShine by China, performs calculations on the order of 10^18 cps (roughly two orders of magnitude more than the human brain) and consumes close to 40 megawatts of continuous power, whereas the human brain performing 10^16 cps consumes only 20 watts of continuous power. Thus, even today, the human brain is extraordinarily energy-efficient compared to the fastest supercomputers.
Future Ahead
These questions had not become mainstream earlier, but with the advent of AI, it is becoming increasingly difficult to ignore them because AI and tech influence everyone.
Coming back to our original point of rogue and misaligned behaviour of AI, it must be noted that this is likely just the beginning and the tip of the iceberg. As AI becomes more efficient in the future, it will become even smarter and it will find even better ways to trick us. It all seems too logical. Humans will continue to put guardrails and better detection mechanisms to avoid catastrophic situations. Basically, humans will have to adapt to the AI and in the future, it will become increasingly difficult.
As the intelligence level of AI climbs higher and higher, the reliance of humans on AI will increase, and a point may come when it will become actually impossible for humans to detect the tricks of the AI. The bridge between human intelligence and AI is brain-machine interfaces, but before that humans will have to be doubly vigilant in making sure that, in its overenthusiasm, AI does not end up harming human interests.
