I have been thinking about a new kind of translation problem: how to turn visual knowledge into something an LLM can actually work with.
Designers are used to reading an image as more than an image. We can see hierarchy, weight, rhythm, proximity, crop, contrast, and intention. We can tell when a logo is too small, when a headline needs room to breathe, or when two elements technically align but do not belong together. Most of that knowledge never arrives as a clean list of instructions. It is held in the reference, in the eye, and in the accumulated experience of making things.
An LLM cannot use that knowledge in the same way. It can look at a reference and say useful things about it, but a finished image is still an output. To make it editable, repeatable, or useful inside a real workflow, the visual needs to become a description of relationships: this is the subject, this is the focal point, this is a background texture, this area is intentionally quiet, this object is meant to feel heavier than the one beside it, and this is the decision that should survive when the layout changes.
A useful comparison: old games
I keep coming back to decompilation. An old game can survive as a compiled thing that runs, but that is not the same as having the source code that explains how it works. Decompilation is the difficult act of recovering enough structure from the finished behavior that people can understand it, edit it, and build it again.
The analogy is not perfect, but it is useful. A screenshot, a flattened graphic, or a finished page is a compiled artifact. It contains the results of thousands of decisions, but not their names. It does not tell you which distance is a system rule and which one is a one-off adjustment. It does not tell you whether a crop is deliberate, whether a shadow carries meaning, or whether a visual pattern should become a reusable component.
So the work is not to ask an LLM to copy pixels. It is to decompile the visual into an editable model of the work: objects, layers, type, rules, constraints, and intent. Then the model can help recompile it into another form, whether that is a web page, a Photoshop file, an Illustrator mark, a presentation, or a new variation of the original design.
That matters because copying the surface and preserving the design are different jobs. A close visual imitation can still lose the actual system. A better translation can change the format while keeping the things that made the source work in the first place.
What feels different with Astra
Working with Astra has made this feel more practical. It is not that the model suddenly has a secret design file inside its head, and I do not think it makes sense to pretend we know exactly how it represents an image. But its spatial understanding changes the conversation.
I can show it a composition and spend less time naming every object before we can start talking about the work. It can track the relationship between a foreground element and its background, identify grouping, notice when a visual is being framed by the page around it, and hold onto the fact that an element has moved rather than simply disappeared. That does not replace judgment. It makes a visual reference more available as working context.
For me, the interesting shift is not simply that the model can describe an image more accurately. It is that spatial understanding makes visual feedback more conversational. I can say that something needs to sit lower, feel more anchored, create a quieter place for text, or stop competing with the thing beside it. Those are incomplete instructions, but they are meaningful design instructions. A model that can keep track of space has a better chance of helping turn them into an actual change.
Adobe is where the translation gets real
This becomes most useful when the model can work with the tools where the design already lives. Adobe applications are full of structured visual knowledge: layers, artboards, selections, masks, paragraph styles, vector paths, color values, frames, exports, and document settings. Those are not just technical details. They are the handles that make a visual idea editable.
Model Context Protocol tools can make those handles available to an LLM in a more concrete way. MCP is not an art director and it is not a replacement for an interface. It is a way to give a model bounded access to information and actions. Instead of describing a document from memory and hoping the model guesses correctly, a tool can inspect what is there, report back on the structure, and make a specific change with a record of what happened.
That is a very different relationship with AI than asking for a new image from scratch. The model can inspect a file, identify the editable pieces, help make a targeted change, and return the work to the same environment where a person can review it. The designer stays close to the artifact instead of throwing it into a black box and hoping it comes back intact.
The important word is bounded. A good tool does not give an LLM an endless pile of controls and expect good taste to emerge. It gives the model a useful vocabulary for the task: read the layer stack, find the text frame, inspect the artboard, change this color, export this version, or leave the document alone and explain what it sees. The scope of the tool shapes the quality of the collaboration.
The human part does not disappear
There is a risk in this comparison. Decompilation sounds objective, as if the right structure is always waiting to be extracted. Visual work is not like that. A flattened image contains ambiguity, and some of the most important decisions are not visible at all: what the client rejected, what had to work on a phone, what was chosen because it could be edited later, and what only exists because there was a deadline.
That is why translation still needs a person. The model can help identify patterns, inventory objects, propose a component structure, and perform the tedious parts of rebuilding. But someone still needs to decide what is signal and what is noise. Someone has to say which relationship is essential, which oddity is part of the character, and which shortcut is going to create trouble later.
I think that is also why good tool use matters more than simply having a capable model. The goal is not to remove the human from the visual process. It is to create a clearer bridge between what a person can see and what the system can act on.
From reference to a shared language
The workflow I want is simple to describe, even if it is hard to make well. Start with the reference. Let the model help turn it into a structured understanding of the visual. Use tools to inspect and change the real working file. Rebuild the output. Then look at it again.
That last step matters. Recompiling an old game does not prove that every recovered decision was right. You run it. You play it. You find what broke. Visual work needs the same loop. The rendered result is the test of whether the translation kept the knowledge that mattered.
This is where I think the next stage of AI-assisted creative work gets interesting. The question is no longer only whether a model can generate an image or write code. It is whether we can build a shared language between visual references, editable files, tool calls, and human judgment. If we can, an image stops being the end of the process. It becomes knowledge that can move.