<?xml version="1.0" encoding="utf-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom">
<channel>
  <title>in Lotusland</title>
  <link>https://in.amagi.dev/</link>
  <description>Casual writing in English by Takayosi Amagi.</description>
  <language>en</language>
  <atom:link href="https://in.amagi.dev/feed.xml" rel="self" type="application/rss+xml"/>
  <lastBuildDate>Wed, 23 Sep 2026 00:00:00 GMT</lastBuildDate>
  <item>
    <title>Jev VJ</title>
    <link>https://in.amagi.dev/posts/jev-vj/</link>
    <guid isPermaLink="true">https://in.amagi.dev/posts/jev-vj/</guid>
    <pubDate>Wed, 23 Sep 2026 00:00:00 GMT</pubDate>
    <description>&lt;p&gt;Like everyone else on the Internet, I also created a Jev demo: a VJ app with Jev.&lt;/p&gt;
&lt;blockquote data-align=&quot;center&quot; class=&quot;twitter-tweet&quot;&gt;&lt;p lang=&quot;en&quot; dir=&quot;ltr&quot;&gt;I made a &lt;a href=&quot;https://x.com/hashtag/Jev?src=hash&amp;amp;ref_src=twsrc%5Etfw&quot;&gt;#Jev&lt;/a&gt; VJ demo.&lt;br/&gt;&lt;br/&gt;Jev listens to what I say, suggests the matching video clips to me, then I manually switch them.&lt;br/&gt;&lt;br/&gt;I wrote a list of clips &amp;amp; effects with short description beforehand, so it can pick the best combination, just like the official &amp;quot;What color is the sky&amp;quot; demo. &lt;a href=&quot;https://t.co/0YVfoa6INP&quot;&gt;https://t.co/0YVfoa6INP&lt;/a&gt; &lt;a href=&quot;https://t.co/4S5SoDgoEJ&quot;&gt;pic.twitter.com/4S5SoDgoEJ&lt;/a&gt;&lt;/p&gt;— 𝘼𝙈𝘼𝙂𝙄 (@amagitakayosi) &lt;a href=&quot;https://x.com/amagitakayosi/status/2102538676375109841?ref_src=twsrc%5Etfw&quot;&gt;September 22, 2026&lt;/a&gt;&lt;/blockquote&gt;

&lt;p&gt;Code: https://github.com/fand/jev-vj&lt;/p&gt;
&lt;p&gt;In this demo, Jev works as a companion.
To the user prompt in the textbox at the bottom, Jev picks the optimal combinations of clip &amp;amp; effect. User switches them manually, so the videos actually play perfectly in sync with the music.&lt;/p&gt;
&lt;p&gt;I'll explain what I've found while building this, and my thoughts on AI-aided visual development.&lt;/p&gt;
&lt;h2&gt;Architecture&lt;/h2&gt;
&lt;p&gt;So like you all, I just stumbled upon the post by Diogo. I found it interesting, especially Jev's ability to play Doom. I'm a graphics nerd and I'm always looking for new ways for VJ-ing, and VJ-ing is kinda similar to playing video games, I thought Jev can lead us to a new VJ experience.&lt;/p&gt;
&lt;p&gt;But soon I hit a wall. Jev's gameplay ability is based on its &lt;em&gt;Choice&lt;/em&gt; model, which excels at finding the optimal answer from the provided &lt;em&gt;choices&lt;/em&gt;. To put this into practice we need to provide the current game state, in other words the state must be evaluable. However in VJ performance the evaluation is our perception, there's no score.&lt;/p&gt;
&lt;p&gt;Also, even with Jev's blazing-fast response we cannot let it switch the videos exactly on the beat. Here in Vancouver, Jev's response time is around 90ms-260ms, which is roughly one or two 16th notes at 120BPM, still too much latency for music / live visuals.&lt;/p&gt;
&lt;p&gt;So I changed the approach; I let Jev just suggest candidates for the next clip. I created a full clip list with metadata like color, bpm, camera movement etc, and included it in the context of the prompt. Jev uses it as the criteria to score the choices and suggest clips in async, then we evaluate the suggested clips and switch them on the beat. The latency doesn't matter anymore.&lt;/p&gt;
&lt;p&gt;&lt;img src=&quot;https://in.amagi.dev/posts/jev-vj/2a449e15-b869-4027-85e1-9f1fc011a026.webp&quot; alt=&quot;Clip suggestion flow diagram&quot;/&gt;&lt;/p&gt;
&lt;p&gt;I let Astra create a demo with this approach. It required a few hours of iteration to stabilize the UI and fix bugs / performance issues, but finally it worked!&lt;/p&gt;
&lt;p&gt;&lt;img src=&quot;https://in.amagi.dev/posts/jev-vj/1de8e726-ee94-4680-b7a1-b13e98da0707.webp&quot; alt=&quot;Screenshot 2026-09-25 at 2.03.06 PM&quot;/&gt;&lt;/p&gt;
&lt;video controls src=&quot;https://in.amagi.dev/posts/jev-vj/cd68e1c0-0d24-4dc8-af30-c28bed25fe81.webm&quot;&gt;&lt;/video&gt;
&lt;h2&gt;Clip Library&lt;/h2&gt;
&lt;p&gt;The biggest problems heres is obviously the library. We have to provide a thorough list of the clips with descriptions concise yet detailed enough to let Jev find the best clips. Unfortunately I couldn't find a way to fully automate this preparation. Here's what I tried:&lt;/p&gt;
&lt;p&gt;First i tried extracting every 10th frames from a video with ffmpeg, then fed them to Astra. It comprehends static images pretty well, but Astra cannot understand the motion inside the frame sequence. I also tried increasing frame frequency but it didn't help.&lt;/p&gt;
&lt;p&gt;I heard Gemini can handle videos pretty well, so I tried it next. At first it looked promising; Gemini gave me a full list of features in the video, like color palette, materials, and the camera movement. However I found that it cannot remember the details when there are many moving objects in the video. One example is this video by &lt;a href=&quot;https://mantissa.xyz/&quot;&gt;Mantissa&lt;/a&gt;; there are thousands of particles moving rapidly and a hard light flicker, but Gemini described this clip &amp;quot;ideal for slow ambient music&amp;quot; just looking at the camera movement speed.&lt;/p&gt;
&lt;p&gt;&lt;img src=&quot;https://in.amagi.dev/posts/jev-vj/c0da3000-728b-486d-b01b-89efbfd1e6dc.webp&quot; alt=&quot;image&quot;/&gt;&lt;/p&gt;
&lt;p&gt;For this demo, I took a semi-manual approach. I let Astra prepare a Markdown template with filenames, then I wrote description for every video using AquaVoice. This was tedious but most flexible as we can edit it easily 😇&lt;/p&gt;
&lt;p&gt;Later I made a library editor, where we can manage those clip data in a spreadsheet. It's still tedious... but I think we need this kind of manual library work for any kind of VJ work, even with completely different AI-aided systems.&lt;/p&gt;
&lt;p&gt;&lt;img src=&quot;https://in.amagi.dev/posts/jev-vj/10073b22-78f7-43ff-8b6b-7884d75b5d36.webp&quot; alt=&quot;image&quot;/&gt;&lt;/p&gt;
&lt;h2&gt;Token limit&lt;/h2&gt;
&lt;p&gt;Once you got a full detailed clip list, it leads to another problem; token limit. With 112 clips I used, the token count was 27k, 84% out of 32 limit. The more clips you add and more detailed descriptions you write, you'll face the severe token limitation... This is a common problem in LLMs though.&lt;/p&gt;
&lt;p&gt;Here's an excerpt from the API request:&lt;/p&gt;
&lt;pre&gt;&lt;code class=&quot;hljs language-json&quot;&gt;&lt;span class=&quot;hljs-punctuation&quot;&gt;{&lt;/span&gt;
  &lt;span class=&quot;hljs-attr&quot;&gt;&amp;quot;model&amp;quot;&lt;/span&gt;&lt;span class=&quot;hljs-punctuation&quot;&gt;:&lt;/span&gt; &lt;span class=&quot;hljs-string&quot;&gt;&amp;quot;jev-latest&amp;quot;&lt;/span&gt;&lt;span class=&quot;hljs-punctuation&quot;&gt;,&lt;/span&gt;
  &lt;span class=&quot;hljs-attr&quot;&gt;&amp;quot;questions&amp;quot;&lt;/span&gt;&lt;span class=&quot;hljs-punctuation&quot;&gt;:&lt;/span&gt; &lt;span class=&quot;hljs-punctuation&quot;&gt;{&lt;/span&gt;
    &lt;span class=&quot;hljs-attr&quot;&gt;&amp;quot;clip&amp;quot;&lt;/span&gt;&lt;span class=&quot;hljs-punctuation&quot;&gt;:&lt;/span&gt; &lt;span class=&quot;hljs-punctuation&quot;&gt;{&lt;/span&gt;
      &lt;span class=&quot;hljs-attr&quot;&gt;&amp;quot;type&amp;quot;&lt;/span&gt;&lt;span class=&quot;hljs-punctuation&quot;&gt;:&lt;/span&gt; &lt;span class=&quot;hljs-string&quot;&gt;&amp;quot;choice&amp;quot;&lt;/span&gt;&lt;span class=&quot;hljs-punctuation&quot;&gt;,&lt;/span&gt;
      &lt;span class=&quot;hljs-attr&quot;&gt;&amp;quot;instructions&amp;quot;&lt;/span&gt;&lt;span class=&quot;hljs-punctuation&quot;&gt;:&lt;/span&gt; &lt;span class=&quot;hljs-string&quot;&gt;&amp;quot;Rank the available footage for a VJ preview deck using state.prompt and state.clips.&amp;quot;&lt;/span&gt;&lt;span class=&quot;hljs-punctuation&quot;&gt;,&lt;/span&gt;
      &lt;span class=&quot;hljs-attr&quot;&gt;&amp;quot;criteria&amp;quot;&lt;/span&gt;&lt;span class=&quot;hljs-punctuation&quot;&gt;:&lt;/span&gt; &lt;span class=&quot;hljs-punctuation&quot;&gt;{&lt;/span&gt;
        &lt;span class=&quot;hljs-attr&quot;&gt;&amp;quot;Opti/Opti6.mov&amp;quot;&lt;/span&gt;&lt;span class=&quot;hljs-punctuation&quot;&gt;:&lt;/span&gt; &lt;span class=&quot;hljs-string&quot;&gt;&amp;quot;Inorganic metal square tunnel with multiple grilles. White light dots travel upward on the grilles while the camera advances at medium speed.&amp;quot;&lt;/span&gt;&lt;span class=&quot;hljs-punctuation&quot;&gt;,&lt;/span&gt;
        ...
      &lt;span class=&quot;hljs-punctuation&quot;&gt;}&lt;/span&gt;
    &lt;span class=&quot;hljs-punctuation&quot;&gt;}&lt;/span&gt;
  &lt;span class=&quot;hljs-punctuation&quot;&gt;}&lt;/span&gt;&lt;span class=&quot;hljs-punctuation&quot;&gt;,&lt;/span&gt;
  &lt;span class=&quot;hljs-attr&quot;&gt;&amp;quot;state&amp;quot;&lt;/span&gt;&lt;span class=&quot;hljs-punctuation&quot;&gt;:&lt;/span&gt; &lt;span class=&quot;hljs-punctuation&quot;&gt;{&lt;/span&gt;
    &lt;span class=&quot;hljs-attr&quot;&gt;&amp;quot;prompt&amp;quot;&lt;/span&gt;&lt;span class=&quot;hljs-punctuation&quot;&gt;:&lt;/span&gt; &lt;span class=&quot;hljs-string&quot;&gt;&amp;quot;white light metallic machine&amp;quot;&lt;/span&gt;&lt;span class=&quot;hljs-punctuation&quot;&gt;,&lt;/span&gt;
    &lt;span class=&quot;hljs-attr&quot;&gt;&amp;quot;action&amp;quot;&lt;/span&gt;&lt;span class=&quot;hljs-punctuation&quot;&gt;:&lt;/span&gt; &lt;span class=&quot;hljs-string&quot;&gt;&amp;quot;candidates&amp;quot;&lt;/span&gt;&lt;span class=&quot;hljs-punctuation&quot;&gt;,&lt;/span&gt;
    &lt;span class=&quot;hljs-attr&quot;&gt;&amp;quot;clips&amp;quot;&lt;/span&gt;&lt;span class=&quot;hljs-punctuation&quot;&gt;:&lt;/span&gt; &lt;span class=&quot;hljs-punctuation&quot;&gt;[&lt;/span&gt;
      &lt;span class=&quot;hljs-comment&quot;&gt;// Detailed description&lt;/span&gt;
      &lt;span class=&quot;hljs-punctuation&quot;&gt;{&lt;/span&gt;
        &lt;span class=&quot;hljs-attr&quot;&gt;&amp;quot;id&amp;quot;&lt;/span&gt;&lt;span class=&quot;hljs-punctuation&quot;&gt;:&lt;/span&gt; &lt;span class=&quot;hljs-string&quot;&gt;&amp;quot;Opti/Opti6.mov&amp;quot;&lt;/span&gt;&lt;span class=&quot;hljs-punctuation&quot;&gt;,&lt;/span&gt;
        &lt;span class=&quot;hljs-attr&quot;&gt;&amp;quot;description&amp;quot;&lt;/span&gt;&lt;span class=&quot;hljs-punctuation&quot;&gt;:&lt;/span&gt; &lt;span class=&quot;hljs-string&quot;&gt;&amp;quot;Inorganic metal square tunnel with multiple grilles. White light dots travel upward on the grilles while the camera advances at medium speed.&amp;quot;&lt;/span&gt;&lt;span class=&quot;hljs-punctuation&quot;&gt;,&lt;/span&gt;
        &lt;span class=&quot;hljs-attr&quot;&gt;&amp;quot;attributes&amp;quot;&lt;/span&gt;&lt;span class=&quot;hljs-punctuation&quot;&gt;:&lt;/span&gt; &lt;span class=&quot;hljs-punctuation&quot;&gt;{&lt;/span&gt;
          &lt;span class=&quot;hljs-attr&quot;&gt;&amp;quot;color.palette&amp;quot;&lt;/span&gt;&lt;span class=&quot;hljs-punctuation&quot;&gt;:&lt;/span&gt; &lt;span class=&quot;hljs-punctuation&quot;&gt;[&lt;/span&gt;&lt;span class=&quot;hljs-string&quot;&gt;&amp;quot;white&amp;quot;&lt;/span&gt;&lt;span class=&quot;hljs-punctuation&quot;&gt;]&lt;/span&gt;&lt;span class=&quot;hljs-punctuation&quot;&gt;,&lt;/span&gt;
          &lt;span class=&quot;hljs-attr&quot;&gt;&amp;quot;material.types&amp;quot;&lt;/span&gt;&lt;span class=&quot;hljs-punctuation&quot;&gt;:&lt;/span&gt; &lt;span class=&quot;hljs-punctuation&quot;&gt;[&lt;/span&gt;&lt;span class=&quot;hljs-string&quot;&gt;&amp;quot;metal&amp;quot;&lt;/span&gt;&lt;span class=&quot;hljs-punctuation&quot;&gt;]&lt;/span&gt;&lt;span class=&quot;hljs-punctuation&quot;&gt;,&lt;/span&gt;
          &lt;span class=&quot;hljs-attr&quot;&gt;&amp;quot;camera.translation&amp;quot;&lt;/span&gt;&lt;span class=&quot;hljs-punctuation&quot;&gt;:&lt;/span&gt; &lt;span class=&quot;hljs-punctuation&quot;&gt;[&lt;/span&gt;&lt;span class=&quot;hljs-string&quot;&gt;&amp;quot;forward_into_scene&amp;quot;&lt;/span&gt;&lt;span class=&quot;hljs-punctuation&quot;&gt;]&lt;/span&gt;&lt;span class=&quot;hljs-punctuation&quot;&gt;,&lt;/span&gt;
          &lt;span class=&quot;hljs-attr&quot;&gt;&amp;quot;camera.translation_speed&amp;quot;&lt;/span&gt;&lt;span class=&quot;hljs-punctuation&quot;&gt;:&lt;/span&gt; &lt;span class=&quot;hljs-string&quot;&gt;&amp;quot;medium&amp;quot;&lt;/span&gt;&lt;span class=&quot;hljs-punctuation&quot;&gt;,&lt;/span&gt;
          ...
        &lt;span class=&quot;hljs-punctuation&quot;&gt;}&lt;/span&gt;&lt;span class=&quot;hljs-punctuation&quot;&gt;,&lt;/span&gt;
        ...
      &lt;span class=&quot;hljs-punctuation&quot;&gt;}&lt;/span&gt;
    &lt;span class=&quot;hljs-punctuation&quot;&gt;]&lt;/span&gt;&lt;span class=&quot;hljs-punctuation&quot;&gt;,&lt;/span&gt;
    &lt;span class=&quot;hljs-attr&quot;&gt;&amp;quot;available_ids&amp;quot;&lt;/span&gt;&lt;span class=&quot;hljs-punctuation&quot;&gt;:&lt;/span&gt; &lt;span class=&quot;hljs-punctuation&quot;&gt;[&lt;/span&gt;...&lt;span class=&quot;hljs-punctuation&quot;&gt;]&lt;/span&gt;&lt;span class=&quot;hljs-punctuation&quot;&gt;,&lt;/span&gt; &lt;span class=&quot;hljs-comment&quot;&gt;// Playable clips, total 111&lt;/span&gt;
    &lt;span class=&quot;hljs-attr&quot;&gt;&amp;quot;candidate_ids&amp;quot;&lt;/span&gt;&lt;span class=&quot;hljs-punctuation&quot;&gt;:&lt;/span&gt; &lt;span class=&quot;hljs-punctuation&quot;&gt;[&lt;/span&gt;...&lt;span class=&quot;hljs-punctuation&quot;&gt;]&lt;/span&gt;&lt;span class=&quot;hljs-punctuation&quot;&gt;,&lt;/span&gt; &lt;span class=&quot;hljs-comment&quot;&gt;// Choosable clips, total 109&lt;/span&gt;
    &lt;span class=&quot;hljs-attr&quot;&gt;&amp;quot;excluded_recent_ids&amp;quot;&lt;/span&gt;&lt;span class=&quot;hljs-punctuation&quot;&gt;:&lt;/span&gt; &lt;span class=&quot;hljs-punctuation&quot;&gt;[&lt;/span&gt;
      &lt;span class=&quot;hljs-string&quot;&gt;&amp;quot;ducky3d/Animation 6.mp4&amp;quot;&lt;/span&gt;&lt;span class=&quot;hljs-punctuation&quot;&gt;,&lt;/span&gt;
      &lt;span class=&quot;hljs-string&quot;&gt;&amp;quot;tatsuyam/aqua_10.mp4&amp;quot;&lt;/span&gt;
    &lt;span class=&quot;hljs-punctuation&quot;&gt;]&lt;/span&gt;&lt;span class=&quot;hljs-punctuation&quot;&gt;,&lt;/span&gt;
    &lt;span class=&quot;hljs-attr&quot;&gt;&amp;quot;current_clip_id&amp;quot;&lt;/span&gt;&lt;span class=&quot;hljs-punctuation&quot;&gt;:&lt;/span&gt; &lt;span class=&quot;hljs-string&quot;&gt;&amp;quot;ducky3d/animation 15.mp4&amp;quot;&lt;/span&gt;&lt;span class=&quot;hljs-punctuation&quot;&gt;,&lt;/span&gt;
    &lt;span class=&quot;hljs-attr&quot;&gt;&amp;quot;current_clip&amp;quot;&lt;/span&gt;&lt;span class=&quot;hljs-punctuation&quot;&gt;:&lt;/span&gt; &lt;span class=&quot;hljs-punctuation&quot;&gt;{&lt;/span&gt; ... &lt;span class=&quot;hljs-punctuation&quot;&gt;}&lt;/span&gt;&lt;span class=&quot;hljs-punctuation&quot;&gt;,&lt;/span&gt; &lt;span class=&quot;hljs-comment&quot;&gt;// Short description of the current clip&lt;/span&gt;
    &lt;span class=&quot;hljs-attr&quot;&gt;&amp;quot;current_effects&amp;quot;&lt;/span&gt;&lt;span class=&quot;hljs-punctuation&quot;&gt;:&lt;/span&gt; &lt;span class=&quot;hljs-punctuation&quot;&gt;[&lt;/span&gt;&lt;span class=&quot;hljs-punctuation&quot;&gt;]&lt;/span&gt;&lt;span class=&quot;hljs-punctuation&quot;&gt;,&lt;/span&gt;
    &lt;span class=&quot;hljs-attr&quot;&gt;&amp;quot;history&amp;quot;&lt;/span&gt;&lt;span class=&quot;hljs-punctuation&quot;&gt;:&lt;/span&gt; &lt;span class=&quot;hljs-punctuation&quot;&gt;[&lt;/span&gt;...&lt;span class=&quot;hljs-punctuation&quot;&gt;]&lt;/span&gt; &lt;span class=&quot;hljs-comment&quot;&gt;// Recently played clip IDs&lt;/span&gt;
  &lt;span class=&quot;hljs-punctuation&quot;&gt;}&lt;/span&gt;
&lt;span class=&quot;hljs-punctuation&quot;&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;To mitigate the token limit, we could structurize the queries and the clip composition. My demo only had 1ch output, but for example, we can sort clips into 2 layers; layer A for overlay clips (2D geometry, patterns, text etc) and layer B for 3D graphics, etc. On request we only send the clip list for the selected layer, will reduce the cotext 50%. We can also choose the category first and let Jev seek videos only in the corresponding sub directory (e.g. &amp;quot;Sci-fi &gt; minimal techno&amp;quot; or &amp;quot;Organic &gt; deep sea&amp;quot;). I haven't tested them yet though.&lt;/p&gt;
&lt;h2&gt;Conclusion&lt;/h2&gt;
&lt;p&gt;I heard this kind of auto-suggestion features are already in DJ apps, but there's no VJ apps providing this feature, AFAIK. Probably that's due to the technical difficulties in video analysis, but Jev or other AI-driven methods can help the situation.&lt;/p&gt;
&lt;p&gt;This was just a rough prototype, but I hope people will find more interesting use cases in live audio/visual world, instead of just generating clips.&lt;/p&gt;</description>
  </item>
  <item>
    <title>Blender as API</title>
    <link>https://in.amagi.dev/posts/blender-as-api/</link>
    <guid isPermaLink="true">https://in.amagi.dev/posts/blender-as-api/</guid>
    <pubDate>Fri, 18 Sep 2026 00:00:00 GMT</pubDate>
    <description>&lt;p&gt;&lt;img src=&quot;https://in.amagi.dev/posts/blender-as-api/7308cb6c-52f2-4631-a63c-2885072fbdc7.webp&quot; alt=&quot;ChatGPT screenshot&quot;/&gt;&lt;/p&gt;
&lt;p&gt;I heard GPT6 Astra has strong capability in 3D related tasks. I saw a tweet of my friend, showing mulitple VJ clips created by Astra with Blender, instead of let GenAIs generate video clips directly:&lt;/p&gt;
&lt;blockquote data-align=&quot;center&quot; class=&quot;twitter-tweet&quot;&gt;&lt;p lang=&quot;ja&quot; dir=&quot;ltr&quot;&gt;AstraがBlenderでVJ素材作るやつ色々研究してた &lt;a href=&quot;https://t.co/urMPamSCj7&quot;&gt;https://t.co/urMPamSCj7&lt;/a&gt; &lt;a href=&quot;https://t.co/1vWqflFP81&quot;&gt;pic.twitter.com/1vWqflFP81&lt;/a&gt;&lt;/p&gt;— Saina (@SainaKey) &lt;a href=&quot;https://x.com/SainaKey/status/2098343985488285869?ref_src=twsrc%5Etfw&quot;&gt;September 11, 2026&lt;/a&gt;&lt;/blockquote&gt;

&lt;p&gt;It makes more sense to me than just generating videos directly with Nano Banana etc; it gives us more control over the entire video generation workflow. Astra generates a Python script as a 3D scene setup, easy to version control. We can edit the materials on Blender, then let Astra reuse it to generate a variety of video clips with a consistent visual taste.&lt;/p&gt;
&lt;figure&gt;&lt;p&gt;&lt;img src=&quot;https://in.amagi.dev/posts/blender-as-api/47709918-1771-463f-9c91-3e6d677555ea.webp&quot; alt=&quot;image&quot;/&gt;
&lt;figcaption&gt;First output&lt;/figcaption&gt;&lt;/p&gt;&lt;/figure&gt;
&lt;p&gt;First I started with a no-brainer prompt: &amp;quot;Create a cool VJ clip on Blender&amp;quot;. Ofc the result was terrible, but at least I could confirm that the Astra + Blender workflow actually does the job.&lt;/p&gt;
&lt;p&gt;Second I gave it some images as reference and said &amp;quot;Create geometric scene with metallic materials like&amp;quot;. Astra made clips kinda okay, but it lacked the sense of color, post-processing and details. However after I threw more requests like &amp;quot;Add bloom, color lights and aberration&amp;quot; it made some progress, it was convincing to me that we can improve the output by telling what pleases our eyes betterf, just like when we do UI dev.&lt;/p&gt;
&lt;p&gt;&lt;img src=&quot;https://in.amagi.dev/posts/blender-as-api/ab8b0d29-70ae-46d8-9359-159ff3b99edf.webp&quot; alt=&quot;Alloy_look_comparison&quot;/&gt;&lt;/p&gt;
&lt;p&gt;Then I remembered Blender 5.3 added dispersion support (rainbow-ish refaction of transparent materials like glass, water etc).&lt;/p&gt;
&lt;p&gt;I asked him to provide me a list of ideas with test render. It was surprising that it made a variety of scenes, though some of them didn't make sense as VJ clip.&lt;/p&gt;
&lt;p&gt;&lt;img src=&quot;https://in.amagi.dev/posts/blender-as-api/97568a75-4409-44f0-a353-7210bdb07c94.webp&quot; alt=&quot;image&quot;/&gt;&lt;/p&gt;
&lt;p&gt;I picked some of them, added some ideas, and let it proceed to full render. Yet they are too lame to be used for VJ directly, they look just a few steps away from what I'd say it's finished.
You know it's AI, the cold bloodless thing, it always try to use dumb simple shapes like square, triangle etc. In the glass flacture sample, all the fractures were strait-cut triangles. It took 20~30 iterations to make the shapes naturalistic, finess the lights and materials, adjust the post-effects.&lt;/p&gt;
&lt;p&gt;The final outputs were like this:&lt;/p&gt;
&lt;video controls src=&quot;https://in.amagi.dev/posts/blender-as-api/ed582022-56d1-47f3-828d-05c95b6fa806.webm&quot;&gt;&lt;/video&gt;
&lt;video controls src=&quot;https://in.amagi.dev/posts/blender-as-api/2534be33-ab07-478f-a007-786554ab1a10.webm&quot;&gt;&lt;/video&gt;
&lt;video controls src=&quot;https://in.amagi.dev/posts/blender-as-api/a4a5d628-9a17-48a7-be75-7a0a0122d4c9.webm&quot;&gt;&lt;/video&gt;
&lt;video controls src=&quot;https://in.amagi.dev/posts/blender-as-api/40813460-3ac4-4247-9c88-85ba62ce2b12.webm&quot;&gt;&lt;/video&gt;
&lt;p&gt;To be accurate, I've seen many folks trying to control Blender with AI since GPT-3.5 era. But thanks to the fast iteration &amp;amp; strong compture-use ability in Astra, it's like Blender is acting like a remote renderer API, kinda more purified &amp;amp; abstracted than working as an MCP server.&lt;/p&gt;
&lt;p&gt;However there's still some quirks; it can't actually judge the visual quality of the output by itself. It knows what is thought to be &amp;quot;cool&amp;quot; as a knowledge, it knows what pleases humans' eyes, but its doesn't feel pleasure. At least, not yet.&lt;/p&gt;
&lt;p&gt;For example, I once tried to generate a particle scene. The sketche generated by GPT Image was good, but it failed to reproduce it in Blender. He understands the base shape, but can't reason how those particles should move, unlike we can easily imagine from the image. The output was just a wobbling sphere, without the flare-like flows in the sketch.&lt;/p&gt;
&lt;p&gt;&lt;img src=&quot;https://in.amagi.dev/posts/blender-as-api/830566e3-a00d-4d03-b6df-d52c9fd107d7.webp&quot; alt=&quot;image&quot;/&gt;&lt;/p&gt;
&lt;p&gt;Once manage to tame Astra completely to let it generate a perfect scene, the bottleneck moves from manipulating Blender to:&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;Explain your demand in format AI can read&lt;/li&gt;
&lt;li&gt;Rendering time&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;&lt;code&gt;1.&lt;/code&gt; is　the problem that we're all facing in software development. But in visual development it could be even more difficult, since it's not straightforward to translate the image in our brain into text.
&lt;code&gt;2.&lt;/code&gt; is a hard problem. It solely depends on the path tracing algorithm and the GPU power. We can use faster renderer like EEVEE or choose game engines like Unity / UE, but if you need high quality outputs, it's an necessarry expense——until AGI discovers a blazing fast pathtracing technique.&lt;/p&gt;</description>
  </item>
  <item>
    <title>You keep rebuilidng blog systems instead of writing blog!</title>
    <link>https://in.amagi.dev/posts/yet-another-blog/</link>
    <guid isPermaLink="true">https://in.amagi.dev/posts/yet-another-blog/</guid>
    <pubDate>Wed, 16 Sep 2026 00:00:00 GMT</pubDate>
    <description>&lt;p&gt;Yes, I know... so this is my new blog, focused on casual writing in English.&lt;/p&gt;
&lt;p&gt;Just like softwares, any blogs can get stale.
&lt;a href=&quot;https://blog.amagi.dev/&quot;&gt;My main blog&lt;/a&gt; is written in Japanese. When I started posting on it, at first only my friends read it so I could write anything. But later I started writing more technical stuffs, the blog's topic got kinda narrowed, and I shifted my feeling and daily whatabout to Twitter. And I changed my job, my métier shifted to Web dev to graphics. I tried to write in the same way as before, but I felt weird friction between what I wanted to write and what the blog had grown into. And life went on, every life shift made my blog more rigid and outdated...&lt;/p&gt;
&lt;p&gt;I wanna change this, esp given the current situation of web &amp;amp; socail media, I need to write my feelings more casually.&lt;/p&gt;
&lt;h2&gt;Architecture&lt;/h2&gt;
&lt;p&gt;When I was learning French last year, I used to write blogs on Hugo. But this time I wanted to make my blog more versatile, like allowing JS / iframe because I'd write technical articles with live demos.&lt;/p&gt;
&lt;figure&gt;&lt;p&gt;&lt;img src=&quot;https://in.amagi.dev/posts/yet-another-blog/3a111d95-df4f-421f-9d7a-e2fdd7d6ef60.webp&quot; alt=&quot;My previous french blog&quot;/&gt;
&lt;figcaption&gt;My previous french blog&lt;/figcaption&gt;&lt;/p&gt;&lt;/figure&gt;
&lt;p&gt;First I asked Claude for possible options for my demands. He suggested Astro, Next and other SSG frameworks, but I didn't wanna make my blog complicated... I don't wanna be bothered by npm update, I just want a static site, not a SSG with flashy new technology.&lt;/p&gt;
&lt;p&gt;Then I let Astra (cuz I had almost unused 1mo subscription) to make up minimal blog setup with Vite. The architecture is like this:&lt;/p&gt;
&lt;p&gt;&lt;img src=&quot;https://in.amagi.dev/posts/yet-another-blog/8324be69-8ca4-47a1-a304-beac91c46d82.webp&quot; alt=&quot;architecture diagram&quot;/&gt;&lt;/p&gt;
&lt;p&gt;The build script &amp;amp; runtime is quite short, just ~1k line of JS.
It also gave me a pretty looking editor page. It has to have GitHub auth so it runs on a Cloudflare Worker, but finally no server is needed.&lt;/p&gt;
&lt;figure&gt;&lt;p&gt;&lt;img src=&quot;https://in.amagi.dev/posts/yet-another-blog/fd68e389-ce8b-4850-a4d7-fab367434f2e.webp&quot; alt=&quot;Editor page screenshot&quot;/&gt;
&lt;figcaption&gt;Editor page running on Cloudflare Worker&lt;/figcaption&gt;&lt;/p&gt;&lt;/figure&gt;
&lt;p&gt;I'll try posting my daily stuff as often as I can, let's see...&lt;/p&gt;</description>
  </item>
</channel>
</rss>
