Trying not to lose my old freebie development topics

3dcheapskate3dcheapskate Posts: 2,732
edited July 18 in The Commons

Since the HiveWire3D forums are closing at the end of this month I'm archiving the forum pages I want using the 'Save Page Now' option on the Wayback Macine's web archive homepage. I'm doing this because the Wayback Machine's normal archiving process doesn't always capture everything.

After running into several problems, which I noted with annoyance in this post, I found a procedure that works for me - so I've added the process I'm using to capture the pages, attachments and all, to a later post below , and deleted all the stuff about the problems I was having as that's no longer relevant..

First and foremost, Electro-harpist posted a link to a 7z backup on the Wayback Macine of seven subforums: Poser, Poser Scripts, Poser Shaders, DAZ Studio, DAZ Scripts, DAZ Shaders and  Photoshop -  Hive Wire Forums Backup June 2026.7z : HiveWire3D users : Free Download, Borrow, and Streaming : Internet Archive - a quick check indicates that it has indeed captured all pages of all topics including all images and all attachments.

Anyway, here are the Wayback Machine links for a few of my topics that I don't want to lose:

And CGBytes 'The Node Knows' topics that I thought had gone forever - they're just mostly gone forever...

The following WM search link shows the URLs of all CGBytes forum pages that have ever been captured ​​https://web.archive.org/web/*/http://www.cgbytes.com/community/forums.aspx*, so if you know yor topics ID number (the ###### in the 'http://www.cgbytes.com:80/community/forums.aspx?g=posts&t=######'; topic URL) you should be able to check whether WM ever captured it. Unfortunately there are very few captures of subsequent pages, which have an extra &p2, &p3, etc on the end of the URL - these are probably specific captures of whole topics by people who knew what they were doing.

Post edited by 3dcheapskate on

Comments

  • kprkpr Posts: 420
    edited July 12

    WBM doesn't seem to save every image on image-heavy pages

    If you need to preserve them intact - minus the attachments, which you'd have to do manually - then you could use the free service on the following link and find _some nice, obliging forum_ where you can attach the resulting pdfs: https://pdfservices.com/html-to-pdf

    Post edited by kpr on
  • TaozTaoz Posts: 10,324
    edited July 13

    There is a free browser plugin called SingleFile, I'm using the FireFox version, but there are versions for several other browsers.  It will save the page in your browser with pictures and everything as a single HTML file, in your Windows Downloads folder.  Results are excellent.  I've been looking for a browser based viewer and organizer for the files but haven't been able to find any so far, so I'm thinking of creating one myself.  

    Post edited by Taoz on
  • TaozTaoz Posts: 10,324
    edited July 13

    BTW, tried one of your links to HiveWire (directly, not via Wayback), and it does seem to get the attachments also, just click on them in the saved HTML file loaded on your browser, to save them to the Downloads folder. HTML page in attached zip. 

    singlefile_attachment.png
    833 x 344 - 46K
    zip
    zip
    Random colour tiles.zip
    3M
    Post edited by Taoz on
  • 3dcheapskate3dcheapskate Posts: 2,732
    edited July 13

    kpr said:

    WBM doesn't seem to save every image on image-heavy pages

    If you need to preserve them intact - minus the attachments, which you'd have to do manually - then you could use the free service on the following link and find _some nice, obliging forum_ where you can attach the resulting pdfs: https://pdfservices.com/html-to-pdf

    Nice idea, but I've always found that PDF versions of most web pages are awful. I tried my browser's print to PDF and had to quickly put on my peril-sensitive sunglasses...

    Post edited by 3dcheapskate on
  • 3dcheapskate3dcheapskate Posts: 2,732

    Taoz said:

    There is a free browser plugin called SingleFile, I'm using the FireFox version, but there are versions for several other browsers.  It will save the page in your browser with pictures and everything as a single HTML file, in your Windows Downloads folder.  Results are excellent.  I've been looking for a browser based viewer and organizer for the files but haven't been able to find any so far, so I'm thinking of creating one myself.  

    I'd forgotten about the saving to a single file - I gave up on that after the old Firefox extension that saved to MHT files ceased to exist. The browser I'm using today has a save to single MHTML file option that seems to work okay, but unlike the MHT you can't just change the extension to ZIP and nose around inside it.

  • jmucchiellojmucchiello Posts: 1,753

    If you right click on a page and say Save As... it will create a subdirectory and include all the HTML, CSS, and images needed to save the page. You can then click the HTML in the directory and it should display correctly in the browser. The directory structure could be zipped up if you want to make the page available to someone else. Not a great solution. But perhaps better than a weird PDF.

  • 3dcheapskate3dcheapskate Posts: 2,732
    edited July 17

    jmucchiello said:

    If you right click on a page and say Save As... it will create a subdirectory and include all the HTML, CSS, and images needed to save the page. You can then click the HTML in the directory and it should display correctly in the browser. The directory structure could be zipped up if you want to make the page available to someone else. Not a great solution. But perhaps better than a weird PDF.

    I used to do that, but I didn't like the HTML file + folder pair that that method produced (reason being that renaming the folder caused the connection to break). I much prefer the save as single file option.

    Regardless of all that, I've now realized that when using the Wayback Machine's single page save ('Save page now' option bottom right here) which saves just the single page at the URL you enter, you can then scroll through the original page, copy the URL of each attachment, and save that attachment using the Wayback Machine's same single page save. Once that's done for all attachments the Wayback Machine's capture of the page will also have captures of all the attachments. It's a bit tedious, and I'd recommend a double check that everything's there after it's all done (which I'm doing and noting in the first post here), but it works. 

    Edit: Also if a page from the Wayback Machine has any broken image link icons it's worth right-click > Open image in new tab (that's Chrome, probably same/similar in other browsers) because the image is usually there, just not served.

    ~ ~ ~

    Here's a method that seems to waork to capture complete archives of HiveWire3D forum pages.

    I've found that the Wayback machine's autocrawl captures of forum topics are rather hit and miss, so for anybody trying to ensure that the WayBack Machine has complete archives of particular forum topics of interest there is a somewhat tedious method that appears to work for the HiveWire3D forums, and you don't need to have a WayBack Machine account. Sorry if it's a bit wordy, but better safe than sorry (thinks RDNA forums, SmithMicro forums, CGBytes 'The Node Know's forum...)

    1. Enter the URL* of the first page of the topic into the 'Save Page Now' option bottom right of the main WayBack Machine web page. When you click the Save Page button you'll be asked if you want to log in and whether you want to save error pages (I don't do either) - it then gets to work - at busy times it may tell you that the capture will be delayed for while (the estimated time they give seems quite accurate). Only that single page is saved, and it also appears to capture any inline images. It shows some 'processing' information and then clearly indicates when the capture has been completed, along with a link to the captured page. If you close the page before the capture is complete you can find the capture by doing a normal wayback machine search (but note that the captures are probably URL-precise*)
    2. While it's working on the capture look through the topic page for any attachments. Right click on each attachment and select 'Copy link address' (even if the attachment is an image I use this rather than 'Copy image address'), then open another Wayback Machine web page in another tab, go to the same 'Save Page Now' option on the main , paste the attachment URL, click the 'Save Page' button and capture the attachment. Note that it only seems to allow you to have two ongoing saves at any time
    3. Repeat step 2 for each attachment. (remember that it only seems to allow you to have two ongoing saves at any time)
    4. Repeat steps 1-3 for each subsequent page of the topic. (remember that it only seems to allow you to have two ongoing saves at any time)
    5. Examine the Wayback Machine capture of the first page of the topic to make sure that all the inline images have been captured. If you see any broken-image-link icons right click on them and select 'Open image in new tab' (or equivalent) and hopefully you'll see the image (he WayBack Machine sometimes appears to have problems serving images in image-heavy pages just after the page has been captured - if not
    6. Click on each attachment in turn in the Wayback Machine capture of the first page - it should open a new Wayback Machine page showing how many captures it has of that attachment (usually just one dated the date you did the capture) and automatically download the attachment, which your browser will probably indicate. I like to actually open the downloaded file to ensure it's hunky-dorey and then delete it, happy in the knowledge that the WayBack Machine has it.**
    7. Repeat 6 for the capture of each page.
    8. If you've done all this but there are still missing inline images/attachments then wait until tomorrow and look at the capture again. Hopefully it's all okay now. If there are still missing images then try step 2 but right-click the images that are missing from the capture and use 'Copy image address' - and remember that it only seems to allow you to have two ongoing saves at any time))


    *If you use a URL that includes a message number at the end, then I think it only connects that page archive to that exact URL, so I prefer to strip off any message number from the end of the URL. For multi-page topics I use the actual page number buttons within the topic to open specific pages and use those URLs
    **Until of course the Wayback Machine curls ups it's tootsies and shuffles off this mortal coil.

    check.jpg
    734 x 543 - 52K
    Post edited by 3dcheapskate on
Sign In or Register to comment.